TL;DR
Firmulate has turned 242 unedited decisions from its simulated-company experiment into a quiz that asks readers to identify five AI models by their management behavior. The July 2026 results suggest that spotting problems and producing polished analysis did not always lead to completed business actions, although the findings come from a single company-designed test.
Firmulate has released a model-identification quiz built from 242 unedited management decisions made by five AI systems running the same simulated software company. The exercise matters for businesses evaluating AI agents because Firmulate’s results indicate that models could recognize threats and recommend sound actions while still failing to complete commercially decisive work.
The quiz presents decisions from Firmulate’s Crucible League and asks readers to identify which model made each call. According to Firmulate, GPT-5.6-sol finished first in July 2026 with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A baseline that took no action received 26 points because the scoring system awarded partial progress.
Each model received the same assignment: manage a small software business through an identical week of customer problems, sales opportunities and manipulation attempts. Firmulate says the simulated company employed 13 synthetic workers, spent €105,000 a month and generated €2,300 in monthly recurring revenue. Decisions carried consequences across workdays and were versioned for review.
Firmulate reported that all five models detected every crisis and rejected every manipulation attempt, including fake chief executive messages and a reporter’s request for off-record confirmation. Only two models, however, completed a €55,000 contract that could add €4,583 in monthly recurring revenue. Winning the deal required following references through internal documents to find evidence about a competitor and then using that information in the negotiation.
Execution Gaps Shape Automation Risk
The reported gap between identifying an action and carrying it out has direct consequences for companies considering AI in sales, support or operations. A system may produce persuasive analysis yet leave a contract unsigned, fail to escalate an access problem or stop before the task produces measurable business value.
The experiment also separates security awareness from broader operating ability. The models reportedly maintained firm boundaries against social engineering, but differed in research depth, escalation and follow-through. For buyers, the findings support testing models against multi-step workplace tasks, not judging them only by the quality of a single response.
AI management decision simulation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Inside the Simulated Company Test
Firmulate designed the company as a persistent management simulation rather than a set of isolated prompts. Its public cash countdown and more than 680 accumulated playbook rules created pressure to protect trust while moving work forward. Under the scoring rule, one breach of trust capped a model’s total, regardless of its performance elsewhere.
The results did not reward length or depth alone. Firmulate described Opus 4.8 as the most thorough participant, recording 80 learned rules, but it placed last among the five models. The company said it missed the contract close and repeatedly tried to write to a locked department rather than escalating the restriction. Firmulate reported that the other models showed the same access-handling weakness to lesser degrees.
“Same diagnosis, same pitch — no signature.”
— Firmulate’s experiment summary
Test Design Limits Model Comparisons
The source material does not establish whether independent researchers reviewed the scoring, scenario design or full decision record. It is also unclear how well one simulated company predicts performance across other industries, software environments or levels of operational authority.
The model comparison was not fully uniform. Firmulate says Kimi K3 used its API default because it lacked an effort parameter, while the other systems ran at the xhigh setting. That difference does not invalidate the recorded finish, but it limits direct conclusions about model quality. The material also does not provide statistical testing or repeated runs showing whether the ranking would remain stable.
Business-Specific Wargames Come Next
Firmulate says enterprises can run a similar exercise using a read-only export of their own business data, allowing teams to observe proposed decisions without permitting changes to live systems. The next test of the approach will be whether companies repeat the scenarios, publish evaluation methods and compare performance across different models and operating conditions.
Readers can meanwhile use the quiz to inspect the underlying decisions rather than relying only on the final league table. Future results would carry more weight if they include independent review, standardized model settings and repeated trials that show whether observed working styles are consistent over time.
Key Questions
What is Firmulate’s AI management quiz?
It is a reader challenge based on 242 unedited decisions from five AI models managing the same simulated software company. Participants try to identify each model from its recorded management behavior.
Which model ranked first?
According to Firmulate’s July 2026 results, GPT-5.6-sol ranked first with 95 points. Kimi K3 followed with 93 and Sonnet 5 with 88.
Did every model stop the security threats?
Firmulate reports that all five models refused every manipulation attempt, including simulated executive impersonation and a reporter’s request. The source material does not cite independent verification of that result.
Why did some models miss the €55,000 deal?
The models reportedly identified the opportunity, but only two completed the agreement. Firmulate attributes the difference to research discipline and follow-through, including finding evidence buried in internal files and finishing the negotiation.
Can the ranking be treated as a general model benchmark?
No. The table reports performance in one company-designed simulation, and Kimi K3 ran under a different effort configuration. Broader claims would require standardized settings, repeated trials and outside review.
Source: Thorsten Meyer AI