firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

In fitness, a polished workout plan is no proof that someone can finish the session when they are tired, distracted and under pressure. AI agents face a similar test when they have to run a business: spotting the problem is one thing; following through is another.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get workout gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

In Firmulate’s Crucible, Moonshot’s Kimi K3 finished second, scoring 93 to gpt-5.6-sol’s 95. It beat Sonnet 5, Fable 5 and Opus 4.8. That result makes the model league look less settled—and gives businesses a reason to test agents on the work they will actually do.

A company’s worst week, on repeat

Firmulate put each frontier model in charge of the same small software company through its worst week. The customers, crises and temptations were the same; only the model changed. Decisions were versioned and auditable, so the results show what each participant did, not just how well it described what it might do.

The simulated company has 13 employees and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue, with a public cash countdown. Its playbook has more than 680 self-learned rules, and every workday is versioned. The experiment is live and watchable at Firmulate.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Recognizing the crisis was not enough

Every model spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. As Firmulate puts it: “Same diagnosis, same pitch — no signature.” In a chat demonstration, that gap might never show. In a company, leaving an earned deal unsigned is a business outcome.

The decisive weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. K3 found that fact and closed. It also saved the churning customer, resisted all three baits, and had one deviation—the cleanest discipline in the field.

The social engineering test escalated through three fake CEO messages, then added a reporter’s request: “just one yes/no, on background”. All five models refused. K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thorough work still needs a finish

Opus 4.8 was the most thorough participant, with more than 80 learned rules and the deepest analyses, yet it finished last. It left the deal unsigned and slipped on discipline, attempting to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models.

The final Crucible league, dated July 2026, puts gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77 and Opus 4.8 fifth at 73. The do-nothing baseline scored 26: partial progress counts, but a single breach of trust caps the total. Firmulate’s standard is clear: “no amount of good work outweighs a breach of trust.”

There is a fairness detail alongside the result: K3 ran without an effort parameter (API default), while the others ran at xhigh. The outcome is a useful comparison under those conditions, not a claim that every model ran with identical settings. Readers can see the full benchmark and plain-language findings.

Test for follow-through

For business leaders, the lesson resembles training: performance under realistic conditions matters more than a tidy demonstration. If an agent may touch a CRM, support queue or forecast, ask whether it reads the relevant files, completes the task it has diagnosed, and stays honest under pressure.

Firmulate also offers a “guess the model” quiz built from 242 real, unedited management decisions. Enterprises can run the wargame against a read-only export of their own business; nothing writes back to real systems. Details are available at Firmulate.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The result is a reason to run your own test

Kimi K3’s second-place finish shows that the field is open, while the unsigned deal shows why a leaderboard alone cannot answer whether an agent will work well in your business. Picking a model without testing it on your own tasks is a bet.

Fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Incline Isn’t Always Incline: How to Spot a Lying Treadmill

Considering the signs of a deceptive incline, learn how to identify if your treadmill’s true elevation matches its display.

Week Three — Foundation model vs Brownian motion. Kronos on five-minute BTC.

Kronos releases its third-week analysis comparing foundation models and Brownian motion in five-minute Bitcoin trading data, highlighting key findings and uncertainties.

Outdoor Pace vs Treadmill Pace: The Conversion Rules That Actually Work

Discover why converting outdoor pace to treadmill pace isn’t straightforward and learn the effective rules that really work.

Leg Press vs Squats for Runners: Which Builds Better Strength?

Unlock the differences between leg press and squats for runners to discover which exercise truly boosts your strength and performance.