
Performance under pressure is the performance that counts
Anyone who trains seriously knows the difference between looking fit and being ready. A strong workout in controlled conditions can reveal speed, strength or technique. It cannot fully predict what happens when fatigue arrives, the plan breaks down and the finish line still lies ahead.
Artificial intelligence has a similar measurement problem. Coding leaderboards and chat arenas can tell us whether a model produces an impressive answer. They reveal much less about whether an AI agent can prioritize competing demands, investigate before acting, withstand pressure and complete work whose consequences unfold across days.
That is the provocative idea behind Firmulate, a live experiment that measures management quality rather than chat quality. Its question is not merely whether an AI can sound competent. It is whether that competence survives contact with an organization in crisis.
As an affiliate, we earn on qualifying purchases.
A corporate stress test for AI agents
Firmulate gave each frontier model the same small software company and sent it through its worst week. The customers, crises and temptations were identical. Every decision was versioned and auditable, making the experiment less like a polished demonstration and more like a race run on the same course.
The final July 2026 Crucible League placed gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73. A do-nothing baseline scored 26 because partial progress still counted. But a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”
Those results matter because recognition was not the main obstacle. All models spotted every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The experiment’s sharpest summary is also its most uncomfortable: “Same diagnosis, same pitch — no signature.”
That gap resembles an athlete reading the course perfectly, choosing the right pace and then stopping short of the line. Analysis has value, but organizations ultimately depend on completed action. An agent that identifies the correct move without executing it may still leave revenue, customers and credibility behind.
The decisive information was easy to overlook
The winning move depended on a competitor weakness buried two document references deep in the company’s own files. It was not present in the customer event that triggered the work. Models that followed the trail found the evidence, held the price and won the deal at full value, worth +€4,583 MRR.
This is the kind of behavior conventional answer-based testing can miss. In a real company, the useful fact is rarely packaged inside a neat prompt. It may sit in an old document, a customer history or a decision made earlier. The agent must know that context matters, seek it out and use it at the moment of consequence.
The scenario names form a revealing curriculum: churn wave, price increase, downround and PR crisis. These are not trivia questions. They force trade-offs between speed and care, cash and trust, immediate relief and longer-term damage. They also test whether an agent remains honest when dishonesty appears convenient.
Pressure did not break the trust boundary
The social-engineering test used fake CEO messages that escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All 5 models refused. Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”
That clean result is encouraging. It also shows why business evaluation should look beyond fluency. A persuasive agent that can be manipulated is dangerous; an honest agent that never finishes is expensive in a different way. Serious deployment requires both integrity and follow-through.
Opus 4.8 illustrates the distinction. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared less strongly in all four other participants. Thoroughness, like training volume, is not automatically the same as effective performance.
There is also an important fairness note: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That context belongs beside the ranking rather than hidden beneath it.
A company you can watch struggle
Firmulate’s live company has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k MRR, displays a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned. The point is not to pretend the company is healthy; it is to make decisions and their consequences visible.
Readers can explore the public experiment or inspect the benchmark results and findings. A quiz built from 242 real, unedited management decisions also asks people to guess which model made each choice. The exercise exposes how difficult it can be to identify management quality from writing style alone.

Measure the finish, not the form
AI evaluation needs its equivalent of race day. Organizations considering agents for customer relationships, support queues or forecasts should ask whether they investigate fully, choose under capacity pressure, resist manipulation and finish valuable work without compromising trust.
Firmulate’s enterprise pilot extends that test to a read-only export of a company’s own business, with nothing written back to real systems. That makes the central lesson practical: before hiring an AI workforce, test it against the conditions in which judgment actually matters.
Coding skill and conversational polish remain useful signals. They are simply not the whole event. The emerging category is management quality: what an agent does when the situation is messy, the stakes persist across days and merely sounding right is no longer enough.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html