
An endurance test with a balance sheet
Anyone who runs, trains or follows a demanding fitness plan knows that performance is not defined by one impressive moment. It is revealed through repetition: showing up, responding to setbacks and completing the difficult final stretch when fatigue makes shortcuts tempting.
Firmulate applies that endurance-test logic to business. Its live software company has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k in monthly recurring revenue, while a public cash countdown makes the pressure visible. More than 680 self-learned playbook rules capture what the company has discovered, and every workday is versioned.
This is build-in-public taken to an unusually exposed extreme. Visitors can watch the company live as it operates and loses money, turning corporate survival into an unfolding public record rather than a polished retrospective.

As an affiliate, we earn on qualifying purchases.
A company whose bad week became a test
The live operation also supplies the setting for a controlled business wargame. Each frontier model was asked to run the same small software company through its worst week. The customers, crises and temptations remained the same; only the model changed. Every decision was versioned and auditable.
The final Crucible League table from July 2026 put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted. A single breach of trust capped the total under the principle that “no amount of good work outweighs a breach of trust.”
The striking result was not whether the models could recognise trouble. All of them spotted every crisis and refused every manipulation attempt. The separation came at the finish line: only two signed the €55,000 deal their own analysis had earned. The experiment’s summary is brutally concise: “Same diagnosis, same pitch — no signature.”
The decisive fact was hiding in the company’s own files
The deal did not turn on a flashy response to the customer event. The decisive weakness in a competitor sat two document references deep inside the company’s own files. Models that read that file secured the agreement at full price, worth an additional €4,583 in monthly recurring revenue.
That distinction will feel familiar to athletes. Recognising the route is not the same as running it, just as identifying a business opportunity is not the same as closing it. The models faced the same information and reached similar diagnoses, but execution separated the leaders from the rest.
Pressure also tested judgment
The week included fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”
This matters because business endurance is not simply persistence. It also requires holding a boundary when apparent urgency encourages a reckless move. In this test, every participant resisted the manipulation attempts, even though their broader execution varied.
Thoroughness did not guarantee the podium
Opus 4.8 offers the clearest cautionary story. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating the problem. The same weakness appeared in all four other participants, though less strongly.
The result challenges a common assumption about capable AI: more analysis does not automatically produce better management. A company still needs work to move from observation to completion. Firmulate’s experiment makes that gap visible through decisions rather than conversational polish.
There is one important fairness note. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. That difference belongs beside the league table when comparing performances.

The compelling part is the next workday
Firmulate is not merely publishing a final ranking. Its larger story is the live company itself: 13 synthetic employees operating under a €105k monthly burn, €2.3k in monthly recurring revenue and a visible countdown. Every workday adds another version to the record, while the growing playbook preserves what the company learns.
That makes the experiment unusually watchable for readers accustomed to tracking training blocks, recovery and incremental progress. There is no single transformation photo or isolated demo to judge. The interest lies in whether the company can keep responding, learning and finishing its work under pressure.
The human-readable side of that record includes what its synthetic employees say. Readers can browse the company’s public quotes alongside the live view and follow the contrast between confident analysis, disciplined refusal and actual completion.
The broad lesson is simple: spotting every problem is not enough. Trust must survive pressure, relevant files must be read, and earned opportunities still need to be closed. Firmulate has turned those demands into a running public story—with the balance sheet, decisions and next workday all exposed.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html