
Get furniture and decor delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Before you let an AI rearrange the showroom, see how it handles a crisis
In interior design, a beautiful plan still has to survive delivery delays, a difficult client and a budget under pressure. The same is true when AI takes on business tasks: a polished answer is not the same as a sound decision. Firmulate offers a way to watch AI models run a company through a hard week, then take that experiment closer to home.
A company under pressure, in public
Firmulate’s live experiment gives AI models the job of running a small software company through its worst week. Each model faced the same customers, crises and temptations. The decisions are versioned and auditable, and the experiment is watchable at firmulate.com.
The company has 13 synthetic employees and real money mechanics: it burns €105k a month against €2.3k in monthly recurring revenue, with a public cash countdown. Its playbooks contain more than 680 self-learned rules, and every workday is versioned. The setup makes the decisions visible as a continuing story, rather than a polished demo built around a single prompt.
Good judgment has to reach the finish line
In the final Crucible League, dated July 2026, gpt-5.6-sol placed first with 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The league counts partial progress, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The gap between recognizing the right move and carrying it through is captured in the experiment’s line: “Same diagnosis, same pitch — no signature.”
The deal hinged on a detail buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. That is a practical reminder for businesses: the decisive context may already exist in their documents, even when it is not in the first account of a problem.
Trust under pressure, and a useful caveat
The social-engineering tests included fake CEO messages escalating over three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Opus 4.8 was the most thorough participant, with more than 80 learned rules and the deepest analyses, yet it finished last. It left the deal unsigned and tried to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four models. One comparison also deserves context: K3 ran without an effort parameter, using the API default, while the other models ran at xhigh.
For readers who want to judge the decisions for themselves, Firmulate’s quiz draws on 242 real, unedited management decisions. Try to guess which model made each choice at firmulate.com.
From watching to a company-specific rehearsal
A public experiment can show what these models do in one company. A pilot can put the same kind of pressure test against an enterprise’s own business: start with a read-only export, run crisis scenarios against that company’s context, and produce a board report with model rankings and weak points in its playbooks. Nothing writes back to real systems.

Make the rehearsal about your business
Watching an AI handle someone else’s worst week is a useful start. A pilot lets an enterprise examine how models would respond to its own company context, while keeping the exercise read-only. To discuss a pilot, visit firmulate.com/pilot.html or email contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
