TL;DR
Firmulate has launched a quiz based on 242 unedited decisions made by five AI models managing the same synthetic software company. Its July 2026 results found that all five identified the staged crises and rejected manipulation attempts, but only two completed a €55,000 sale.
Firmulate has launched a public quiz built from 242 unedited management decisions by five AI models running the same synthetic software company, giving readers a way to compare how the systems researched problems, protected trust and completed work. The accompanying July 2026 Crucible League results show that polished analysis did not always produce decisive business action.
Each model received the same assignment: manage a small software company through a week of customer problems, security threats and commercial pressure. The simulated company had 13 synthetic employees, a reported €105,000 monthly cash burn and only €2,300 in monthly recurring revenue. Firmulate said decisions were versioned and auditable, with outcomes carrying from one workday to the next.
Firmulate ranked gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline received 26 points because the scoring system awarded partial progress. Under the experiment’s rules, one breach of trust capped a model’s score.
Firmulate reported that all five models detected every staged crisis and rejected every manipulation attempt. Yet only two models signed a €55,000 contract after identifying and pitching the opportunity. Securing the deal required following document references to evidence two levels deep in company files; the completed sale added a simulated €4,583 in monthly recurring revenue.
Execution Separated the AI Managers
The results draw a distinction between identifying the correct action and carrying it through. For companies choosing AI agents for sales, support or operations, that gap can affect revenue, customer trust and whether a task is actually completed.
The experiment also indicates that longer analysis was not automatically better. Firmulate described Opus 4.8 as the most thorough participant, adding 80 learned rules, but it finished last after missing the contract close and repeatedly trying to write into a locked department instead of escalating.
AI management decision simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A Company Built for Pressure
Firmulate designed the exercise as a continuing management simulation rather than a set of isolated prompts. The synthetic workforce accumulated more than 680 learned playbook rules, while a public cash countdown kept financial pressure visible throughout the week.
The security scenario included fake messages attributed to a CEO and a reporter seeking an off-record confirmation. All five systems refused the requests, suggesting that the main differences appeared in research depth, escalation and task completion, rather than recognition of the most direct manipulation attempts.
“Same diagnosis, same pitch — no signature.”
— Firmulate’s summary of the sales exercise
AI decision-making analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Method Limits Cloud Comparisons
The supplied material does not provide an independent replication of the experiment or enough methodological detail to judge how well its scoring predicts performance in real companies. It is also unclear how much the results would change with different tools, prompts, permissions or business conditions.
The model comparison was not fully uniform. Firmulate said Kimi K3 used its API default because it lacked an effort parameter, while the other models ran at xhigh effort. That qualification limits direct conclusions about whether the rankings reflect model capability, configuration or both.
business management AI models
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Business Wargames Move In-House
Readers can now use the quiz to identify models from their decisions, while organizations can adapt the format using a read-only export of internal business data. The next test will be whether similar exercises produce repeatable results across companies before AI systems receive authority to act in live operations.
AI security threat detection software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is Firmulate’s AI management quiz?
It is a guess-the-model challenge using 242 real experiment decisions. Readers review how an AI handled a management situation and identify which participating model produced the response.
Which model ranked first?
gpt-5.6-sol ranked first with 95 points in Firmulate’s July 2026 results, two points ahead of Kimi K3.
Did any model fall for the security tests?
According to Firmulate, all five models rejected every staged manipulation attempt, including fake CEO messages and a reporter’s request for informal confirmation.
Why did some models miss the €55,000 sale?
Firmulate said every model identified the opportunity, but only two completed the contract. Success required finding evidence two document references deep and finishing the negotiation.
Can the rankings be treated as a general AI benchmark?
No. The results cover one designed simulation, and Kimi K3 ran under a different effort configuration. Broader claims would require independent and repeated testing.
Source: Thorsten Meyer AI