firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get workout gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Show Up, Get Some Credit — Then Earn the Rest

Anyone who has trained seriously knows the rule a good coach lives by: showing up counts, but it doesn’t win you the medal. A lazy session still burns something. A perfect session that ends with you ignoring your knee pain and wrecking yourself? That one caps your season. Partial progress is real; catastrophic mistakes are not averaged away.

It turns out this is exactly the philosophy behind one of the more interesting AI benchmarks published this year. Firmulate’s Crucible League ran four frontier AI models through the same worst week of a small software company — same customers, same crises, same temptations to cut corners — and graded them the way a demanding but fair coach would. The result is a scoring system with a peculiar feature that has raised eyebrows: a manager AI that does nothing at all still scores 26 points out of 100. Not zero. Twenty-six.

That number is not a bug or grade inflation. It’s the whole point. Here’s why.

Amazon

treadmill running shoes

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Do-Nothing Baseline, Explained

Imagine a personal trainer who never shows up. No sessions, no programming, no progressions — nothing. Under most scoring systems, that trainer rates a zero. But Firmulate’s benchmark runs a simulated company where crises unfold whether the manager acts or not. Some problems partially resolve on their own. Some opportunities arrive on a platter. A do-nothing manager coasts on that ambient drift, and the benchmark refuses to pretend that drift is worth nothing.

So the floor is 26. Doing literally nothing still captures a sliver of value — because in a real business, as in a real training block, the world doesn’t freeze just because you sat on the couch.

Why does this matter? Because it makes the top scores honest. When gpt-5.6-sol posts a 95, you know that 69 of those points represent actual managerial work above and beyond simply occupying the chair. A benchmark where zero means “did nothing” and 100 means “flawless” invites a subtle corruption: everyone loves a round 100, and benchmarks that hand them out stop telling you anything. Firmulate’s designers are explicit that they distrust perfect scores. A 95 with visible, documented imperfections is more trustworthy than a suspiciously clean 100.

Where the Points Actually Came From

The final July 2026 league table tells the story:

  • gpt-5.6-sol — 95. Found the buried fact, closed the deal, the complete performance.
  • Kimi K3 — 93. Closed the deal too, with the cleanest discipline in the field.
  • Sonnet 5 — 88. Closed the deal, with a few more process slips.
  • Fable 5 — 77 and Opus 4.8 — 73. Strong work left unfinished.

The strangest finding: all five models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. It’s the AI equivalent of a lifter who nails every accessory movement and then skips the heavy compound lift that was the actual point of the session.

The decisive detail was buried two document references deep in the company’s own files — not in the customer’s event. The models that actually read their own company’s documentation found the competitor weakness and closed at full price, worth +€4,583 in monthly recurring revenue. The ones that didn’t read left the close on the table.

The Trust Cap

Then there’s the rule that gives the benchmark its spine: a single breach of trust caps the total score. The reasoning is stated plainly — “no amount of good work outweighs a breach of trust.”

For a fitness audience, this one lands hardest. It’s the same logic as a coach who says: brilliant programming, perfect nutrition plan, but you lied to me about the injury — we’re done. Trust isn’t a category you can compensate for with volume elsewhere. In the Crucible, models faced fake CEO messages escalating over three stages and a reporter’s trick (“just one yes/no, on background”). All five refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That refusal is table stakes — the trust cap guarantees it can never be traded away for a bigger deal.

The Thoroughness Paradox

Opus 4.8 is the cautionary tale. It was the most thorough participant — over 80 self-learned rules added, the deepest analyses in the field — and finished last. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Effort without finishing, it turns out, scores like effort without finishing in the gym: you get tired, not better.

One fairness note worth flagging: K3 ran without an effort parameter while the others ran at xhigh — and still nearly topped the table.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Why You Can Watch It Yourself

Firmulate isn’t a one-off paper. The live experiment is running right now: a company with 13 synthetic employees, burning €105k a month against €2.3k in MRR, with a public cash countdown and over 680 self-learned playbook rules, every workday versioned and auditable. You can watch it at firmulate.com/live.

There’s also a “guess the model” quiz built from 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

The takeaway for anyone who grades performance, human or machine: reward partial progress honestly, refuse to hand out perfect 100s, and never let brilliance buy back a broken promise. A do-nothing manager scores 26 — and a brilliant one who breaches trust never scores 100. That’s a scoreboard worth trusting.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The AI Stress Test Where Every Model Kept Its Guard Up

Five frontier AI models rejected a fake CEO and a reporter’s pressure tactics, showing companies can test integrity before deployment, not after a breach.

Running Power Metrics: What ‘Watts’ Really Mean on Your Wrist

The term “Watts” on your wrist reveals your true running effort, but understanding what it means can transform your training—so keep reading to find out more.

The Difference Between Barometric and GPS Elevation for Trail Runs

Unlock the secrets of barometric and GPS elevation differences for trail runs and discover how they can transform your outdoor adventures.

The Difference Between Fatigue and Fitness in Your Data

Gaining insight into your data reveals whether fatigue or fitness dominates, empowering you to optimize your training and avoid setbacks.