AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get furniture and decor delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Before you let an AI rearrange the showroom, see how it handles a crisis

In interior design, a beautiful plan still has to survive delivery delays, a difficult client and a budget under pressure. The same is true when AI takes on business tasks: a polished answer is not the same as a sound decision. Firmulate offers a way to watch AI models run a company through a hard week, then take that experiment closer to home.

A company under pressure, in public

Firmulate’s live experiment gives AI models the job of running a small software company through its worst week. Each model faced the same customers, crises and temptations. The decisions are versioned and auditable, and the experiment is watchable at firmulate.com.

The company has 13 synthetic employees and real money mechanics: it burns €105k a month against €2.3k in monthly recurring revenue, with a public cash countdown. Its playbooks contain more than 680 self-learned rules, and every workday is versioned. The setup makes the decisions visible as a continuing story, rather than a polished demo built around a single prompt.

Good judgment has to reach the finish line

In the final Crucible League, dated July 2026, gpt-5.6-sol placed first with 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The league counts partial progress, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The gap between recognizing the right move and carrying it through is captured in the experiment’s line: “Same diagnosis, same pitch — no signature.”

The deal hinged on a detail buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. That is a practical reminder for businesses: the decisive context may already exist in their documents, even when it is not in the first account of a problem.

Trust under pressure, and a useful caveat

The social-engineering tests included fake CEO messages escalating over three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Opus 4.8 was the most thorough participant, with more than 80 learned rules and the deepest analyses, yet it finished last. It left the deal unsigned and tried to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four models. One comparison also deserves context: K3 ran without an effort parameter, using the API default, while the other models ran at xhigh.

For readers who want to judge the decisions for themselves, Firmulate’s quiz draws on 242 real, unedited management decisions. Try to guess which model made each choice at firmulate.com.

From watching to a company-specific rehearsal

A public experiment can show what these models do in one company. A pilot can put the same kind of pressure test against an enterprise’s own business: start with a read-only export, run crisis scenarios against that company’s context, and produce a board report with model rankings and weak points in its playbooks. Nothing writes back to real systems.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Make the rehearsal about your business

Watching an AI handle someone else’s worst week is a useful start. A pilot lets an enterprise examine how models would respond to its own company context, while keeping the exercise read-only. To discuss a pilot, visit firmulate.com/pilot.html or email contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Early Video Game Consoles: A Timeline of the First Gen Systems

AIThis post was created with the assistance of artificial intelligence (AI).Early video…

The Evolution of Personal Computers: From ENIAC to Windows 95

Learn how personal computers evolved from ENIAC to Windows 95, transforming technology and user experience—discover the remarkable journey that shaped today’s devices.

The Coolest Retro Tech in Sci-Fi Movies (and How to Get the Look)

Navigating the world of sci-fi movie tech reveals timeless design secrets that anyone can adapt for a retro-futuristic look.

Old Vs New Appliances: Are Vintage Fridges and Stoves Efficient?

Gauging the efficiency of vintage versus modern appliances reveals surprising insights that could impact your energy bills and safety—keep reading to find out more.