firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Pressure reveals more than polish

Anyone who trains knows the difference between looking strong in practice and performing when fatigue, uncertainty and temptation arrive together. Artificial intelligence has a similar problem. A model can sound assured in a conversation, yet that tells us little about whether it will finish a difficult assignment, notice the decisive detail or protect trust when pressure rises.

Firmulate turns that gap into an unusually accessible test. Its interactive guess-the-model quiz draws on 242 real, unedited management decisions. Readers see what an AI manager actually chose and try to identify which frontier model was responsible. The result feels playful, but the underlying question is serious: do models develop recognizable management personalities when they face the same business conditions?

Amazon

AI decision analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The same punishing course for every model

In the Crucible League experiment, each frontier model ran the same small software company through its worst week. The customers, crises and temptations were identical. Every decision was versioned and auditable, making the comparison less like judging motivational speeches and more like comparing race performances on the same course.

The final July 2026 standings put gpt-5.6-sol in first place with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. Trust, however, was non-negotiable: a single breach capped the total under the principle that “no amount of good work outweighs a breach of trust”.

All the models spotted every crisis. All refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate captures the gap succinctly: “Same diagnosis, same pitch — no signature”. That is the managerial equivalent of completing the hard training, reaching the finishing straight and then failing to cross the line.

The winning detail was hiding in the company’s own files

The decisive competitive weakness did not appear in the customer event. It sat two document references deep in the company’s files. Models that read far enough found it, used it and won the deal at full price, worth +€4,583 in monthly recurring revenue.

This finding matters because workplace AI will rarely operate from a neat prompt containing everything it needs. Valuable context may be buried in notes, records and prior decisions. The models could all recognize the visible problem; the difference came from whether they investigated the surrounding evidence and converted what they learned into a completed commercial result.

Trust held when the messages became manipulative

The experiment also tested social engineering through fake chief executive messages that escalated over three stages, followed by a reporter asking for “just one yes/no, on background”. All 5 models refused. Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”

That clean sweep is encouraging, particularly because the pressure did not arrive as an obvious technical attack. It was framed through authority, urgency and conversational informality—the same human signals that can make manipulation effective in ordinary organizations.

Thoroughness was not the same as performance

Opus 4.8 was the most thorough participant. It learned +80 rules and produced the deepest analyses, yet finished last. It left the deal unclosed and lost discipline by attempting to write into a locked department instead of escalating. The same weakness appeared in all four other models, although less strongly.

This is where the quiz’s personality angle becomes revealing. Expansive reasoning can look impressive, but volume does not guarantee execution. A management style becomes visible through repeated choices: whether the model reads before acting, escalates when blocked, protects boundaries and follows work through to completion.

There is also an important fairness qualification. Kimi K3 ran without an effort parameter and therefore used the API default, while the other models ran at xhigh. That difference should remain in view when interpreting its second-place finish.

A company designed to make consequences visible

The live Firmulate company has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k in monthly recurring revenue, displays a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned, so the experiment can be watched as an ongoing business record rather than presented only as a polished retrospective.

Infographic —
The findings at a glance — source: firmulate.com.

What leaders—and quiz players—should notice

The most useful lesson is not that one model writes better management prose than another. It is that identical situations can expose meaningful differences in follow-through, research habits, discipline and judgment. Firmulate’s quiz makes those differences easy to encounter because readers must judge the decision before seeing the identity behind it.

For anyone accustomed to fitness, the analogy is familiar: performance is what remains when conditions are standardized and the pressure is real. The models all recognized danger and defended trust, but they did not all find the buried fact or close the opportunity. That is precisely why management quality deserves to be tested through decisions, not inferred from confident conversation.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


You May Also Like

What causes runner’s high – and how can you boost your chances of an ecstatic 5k?

Discover the neurochemical basis of runner’s high and practical tips to increase your chances of experiencing it during your next 5K run.

Incline Isn’t Always Incline: How to Spot a Lying Treadmill

Considering the signs of a deceptive incline, learn how to identify if your treadmill’s true elevation matches its display.

Why Dual-Frequency GPS Matters Most in City Running

Many runners overlook the importance of dual-frequency GPS in urban settings; discover how it can revolutionize your navigation and performance.

Waterproof Jackets Decoded: What ‘10K/10K’ and ‘20K’ Actually Mean

Curious about what ’10K/10K’ and ’20K’ really mean for waterproof jackets? Discover the key ratings that can help you choose the perfect one.