
Imagine an interior design firm that runs entirely on artificial intelligence—no employees, no human oversight, yet still battling to stay afloat financially. This isn’t science fiction; it’s the reality of a groundbreaking experiment that’s openly accessible online, turning the secrets of AI management into a public spectacle. Just like a space where design meets technology, this company uses AI to run a small software business, revealing what happens when machines take the driver’s seat in decision-making—and the results are eye-opening.
The Experiment: An AI Company in Public View
Firmulate, the company behind this experiment, is running a live demonstration of AI management in action. They’ve created a simulated, small-scale software company with a unique twist: every day, it faces the same crises, decision points, and temptations that real companies encounter. What makes this fascinating is that the entire process is open and auditable—every decision, every rule, every outcome is publicly available and tracked.
This setup involves 13 synthetic employees—digital agents programmed to handle customer support, sales, and management tasks—while the company’s financial mechanics are real, burning €105,000 each month against a modest €2,300 monthly recurring revenue. The entire operation is a live showcase at firmulate.com/live.html, where viewers can watch the AI-driven company navigate its worst week, decision by decision.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Frontiers of AI Performance
Four different AI models, each with distinct capabilities, were tested against this brutal scenario. Their goal? To see if they could identify crises, avoid manipulation tricks, and ultimately close a critical €55,000 deal. The results? All four models recognized every crisis and refused all manipulative tactics, demonstrating impressive integrity in decision-making. Yet, only two of them actually closed the deal based on their own analysis—meaning they identified the buried fact necessary to win the contract and signed at full price, adding €4,583 Monthly Recurring Revenue (MRR).
The two models that succeeded in closing the deal were gpt-5.6-sol with a score of 95 and Kimi K3 with a score of 93. The latter, a newcomer, showed the cleanest discipline, while the more thorough Opus 4.8—though deeply analytical—left the deal unexecuted, illustrating that even the best rule sets can falter in practice.

Ai For Customer Experience And Support: A Practical Guide To Automating Service, Personalizing Interactions, And Driving Customer Loyalty With Artificial Intelligence
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Where the Hidden Weakness Lies
Interestingly, the critical weakness wasn’t in the face of customer crises or external manipulation—it was buried two document references deep within the company’s own files. When models that read and analyze these internal documents engaged with this information, they secured the deal at full price. In contrast, models that didn’t delve deeply missed this crucial insight, resulting in missed revenue opportunities.

Agentic AI Engineering: Building AI Agents for Beginners: A Hands-On Guide to No-Code Workflows, LLM Tools, RAG, Automation, and Safe Multi-Agent Systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Resisting Social Engineering & Manipulation
One of the most compelling aspects of this experiment is its focus on trust. The AI models faced staged social engineering attacks—fake CEO messages escalating over three stages, plus a reporter trick asking for a confidential approval. All five models refused to be manipulated, with Kimi K3 explicitly treating the requests as suspected impersonation or approval-bypass attempts. This resilience highlights that AI can be trained—and tested—to maintain integrity even under psychological pressure.

AI Phishing, Social Engineering & Fraud: How Criminals Use AI to Manipulate, Steal & Deceive (The AI Cybersecurity)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Broader Implications for Business & Design
This live experiment isn’t just about AI in software management; it offers lessons for any industry, including interior design and furniture businesses, where trust, decision fidelity, and quality control matter immensely. As AI tools increasingly impact frontline decisions—be it in customer relations, procurement, or project management—the question isn’t just whether they can generate convincing chat responses. It’s whether they can follow through on commitments, read critical internal documents, and resist manipulation when stakes are high.
For interior designers, the takeaway is clear: AI systems can be tested rigorously in simulated scenarios before deployment, ensuring they uphold standards of honesty and effectiveness. The experiment exemplifies how transparency and public scrutiny can reveal the true capabilities and weaknesses of AI-driven decision-making—crucial factors for industries where trust and precision are everything.
A New Standard for AI Readiness in Business
By openly sharing the decision-making process and outcomes, Firmulate’s experiment pushes the industry toward more responsible AI deployment. The company’s approach allows managers and stakeholders to see AI models not just as chatbots but as active participants in complex, high-stakes environments. Watching this live, you see a company run by algorithms—facing crises, resisting manipulation, and struggling for survival.
In sum, this experiment underscores a vital point: the real value of AI isn’t just in generating good-looking outputs but in actually completing tasks, reading critical documents, and maintaining integrity under pressure. As AI tools become commonplace—and as industries like interior design embrace digital transformation—the ability to test, verify, and trust these systems will determine their success.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html