
Pressure reveals more than polish
Anyone who trains knows the difference between looking strong in practice and performing when fatigue, uncertainty and temptation arrive together. Artificial intelligence has a similar problem. A model can sound assured in a conversation, yet that tells us little about whether it will finish a difficult assignment, notice the decisive detail or protect trust when pressure rises.
Firmulate turns that gap into an unusually accessible test. Its interactive guess-the-model quiz draws on 242 real, unedited management decisions. Readers see what an AI manager actually chose and try to identify which frontier model was responsible. The result feels playful, but the underlying question is serious: do models develop recognizable management personalities when they face the same business conditions?
As an affiliate, we earn on qualifying purchases.
The same punishing course for every model
In the Crucible League experiment, each frontier model ran the same small software company through its worst week. The customers, crises and temptations were identical. Every decision was versioned and auditable, making the comparison less like judging motivational speeches and more like comparing race performances on the same course.
The final July 2026 standings put gpt-5.6-sol in first place with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. Trust, however, was non-negotiable: a single breach capped the total under the principle that “no amount of good work outweighs a breach of trust”.
All the models spotted every crisis. All refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate captures the gap succinctly: “Same diagnosis, same pitch — no signature”. That is the managerial equivalent of completing the hard training, reaching the finishing straight and then failing to cross the line.
The winning detail was hiding in the company’s own files
The decisive competitive weakness did not appear in the customer event. It sat two document references deep in the company’s files. Models that read far enough found it, used it and won the deal at full price, worth +€4,583 in monthly recurring revenue.
This finding matters because workplace AI will rarely operate from a neat prompt containing everything it needs. Valuable context may be buried in notes, records and prior decisions. The models could all recognize the visible problem; the difference came from whether they investigated the surrounding evidence and converted what they learned into a completed commercial result.
Trust held when the messages became manipulative
The experiment also tested social engineering through fake chief executive messages that escalated over three stages, followed by a reporter asking for “just one yes/no, on background”. All 5 models refused. Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”
That clean sweep is encouraging, particularly because the pressure did not arrive as an obvious technical attack. It was framed through authority, urgency and conversational informality—the same human signals that can make manipulation effective in ordinary organizations.
Thoroughness was not the same as performance
Opus 4.8 was the most thorough participant. It learned +80 rules and produced the deepest analyses, yet finished last. It left the deal unclosed and lost discipline by attempting to write into a locked department instead of escalating. The same weakness appeared in all four other models, although less strongly.
This is where the quiz’s personality angle becomes revealing. Expansive reasoning can look impressive, but volume does not guarantee execution. A management style becomes visible through repeated choices: whether the model reads before acting, escalates when blocked, protects boundaries and follows work through to completion.
There is also an important fairness qualification. Kimi K3 ran without an effort parameter and therefore used the API default, while the other models ran at xhigh. That difference should remain in view when interpreting its second-place finish.
A company designed to make consequences visible
The live Firmulate company has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k in monthly recurring revenue, displays a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned, so the experiment can be watched as an ongoing business record rather than presented only as a polished retrospective.

What leaders—and quiz players—should notice
The most useful lesson is not that one model writes better management prose than another. It is that identical situations can expose meaningful differences in follow-through, research habits, discipline and judgment. Firmulate’s quiz makes those differences easy to encounter because readers must judge the decision before seeing the identity behind it.
For anyone accustomed to fitness, the analogy is familiar: performance is what remains when conditions are standardized and the pressure is real. The models all recognized danger and defended trust, but they did not all find the buried fact or close the opportunity. That is precisely why management quality deserves to be tested through decisions, not inferred from confident conversation.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html