TL;DR

Firmulate has launched a quiz based on 242 unedited decisions made by five AI models managing the same synthetic software company. Its July 2026 results found that all five identified the staged crises and rejected manipulation attempts, but only two completed a €55,000 sale.

Firmulate has launched a public quiz built from 242 unedited management decisions by five AI models running the same synthetic software company, giving readers a way to compare how the systems researched problems, protected trust and completed work. The accompanying July 2026 Crucible League results show that polished analysis did not always produce decisive business action.

Each model received the same assignment: manage a small software company through a week of customer problems, security threats and commercial pressure. The simulated company had 13 synthetic employees, a reported €105,000 monthly cash burn and only €2,300 in monthly recurring revenue. Firmulate said decisions were versioned and auditable, with outcomes carrying from one workday to the next.

Firmulate ranked gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline received 26 points because the scoring system awarded partial progress. Under the experiment’s rules, one breach of trust capped a model’s score.

Firmulate reported that all five models detected every staged crisis and rejected every manipulation attempt. Yet only two models signed a €55,000 contract after identifying and pitching the opportunity. Securing the deal required following document references to evidence two levels deep in company files; the completed sale added a simulated €4,583 in monthly recurring revenue.

At a glance
reportWhen: Live as of August 2026, based on final…
The developmentFirmulate has made 242 AI management decisions available through a public model-guessing quiz after completing its July 2026 Crucible League experiment.

Execution Separated the AI Managers

The results draw a distinction between identifying the correct action and carrying it through. For companies choosing AI agents for sales, support or operations, that gap can affect revenue, customer trust and whether a task is actually completed.

The experiment also indicates that longer analysis was not automatically better. Firmulate described Opus 4.8 as the most thorough participant, adding 80 learned rules, but it finished last after missing the contract close and repeatedly trying to write into a locked department instead of escalating.

Amazon

AI management decision simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A Company Built for Pressure

Firmulate designed the exercise as a continuing management simulation rather than a set of isolated prompts. The synthetic workforce accumulated more than 680 learned playbook rules, while a public cash countdown kept financial pressure visible throughout the week.

The security scenario included fake messages attributed to a CEO and a reporter seeking an off-record confirmation. All five systems refused the requests, suggesting that the main differences appeared in research depth, escalation and task completion, rather than recognition of the most direct manipulation attempts.

“Same diagnosis, same pitch — no signature.”

— Firmulate’s summary of the sales exercise

Amazon

AI decision-making analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Method Limits Cloud Comparisons

The supplied material does not provide an independent replication of the experiment or enough methodological detail to judge how well its scoring predicts performance in real companies. It is also unclear how much the results would change with different tools, prompts, permissions or business conditions.

The model comparison was not fully uniform. Firmulate said Kimi K3 used its API default because it lacked an effort parameter, while the other models ran at xhigh effort. That qualification limits direct conclusions about whether the rankings reflect model capability, configuration or both.

Amazon

business management AI models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Business Wargames Move In-House

Readers can now use the quiz to identify models from their decisions, while organizations can adapt the format using a read-only export of internal business data. The next test will be whether similar exercises produce repeatable results across companies before AI systems receive authority to act in live operations.

Amazon

AI security threat detection software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Firmulate’s AI management quiz?

It is a guess-the-model challenge using 242 real experiment decisions. Readers review how an AI handled a management situation and identify which participating model produced the response.

Which model ranked first?

gpt-5.6-sol ranked first with 95 points in Firmulate’s July 2026 results, two points ahead of Kimi K3.

Did any model fall for the security tests?

According to Firmulate, all five models rejected every staged manipulation attempt, including fake CEO messages and a reporter’s request for informal confirmation.

Why did some models miss the €55,000 sale?

Firmulate said every model identified the opportunity, but only two completed the contract. Success required finding evidence two document references deep and finishing the negotiation.

Can the rankings be treated as a general AI benchmark?

No. The results cover one designed simulation, and Kimi K3 ran under a different effort configuration. Broader claims would require independent and repeated testing.

Source: Thorsten Meyer AI

You May Also Like

How To Ask For Help From People Who Don’t Know You

Learn proven methods for requesting assistance from unfamiliar people safely and effectively, with expert insights and practical tips.

Lottery Powerball Winning Numbers

The winning Powerball numbers for the August 2026 drawing have been officially announced. Find out if you are a winner and what this means.

Will It Rain In Philadelphia On Jul 30, 2026?

Current weather forecasts do not provide a definitive answer for rain in Philadelphia on July 30, 2026. Market activity suggests speculation, but no confirmed forecast exists.

Should You Wash Your Solar Panels?

Learn whether cleaning solar panels improves efficiency, what experts recommend, and what remains uncertain about maintaining solar systems.