firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

In fitness, a polished workout plan is no proof that someone can finish the session when they are tired, distracted and under pressure. AI agents face a similar test when they have to run a business: spotting the problem is one thing; following through is another.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get workout gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

In Firmulate’s Crucible, Moonshot’s Kimi K3 finished second, scoring 93 to gpt-5.6-sol’s 95. It beat Sonnet 5, Fable 5 and Opus 4.8. That result makes the model league look less settled—and gives businesses a reason to test agents on the work they will actually do.

A company’s worst week, on repeat

Firmulate put each frontier model in charge of the same small software company through its worst week. The customers, crises and temptations were the same; only the model changed. Decisions were versioned and auditable, so the results show what each participant did, not just how well it described what it might do.

The simulated company has 13 employees and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue, with a public cash countdown. Its playbook has more than 680 self-learned rules, and every workday is versioned. The experiment is live and watchable at Firmulate.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Recognizing the crisis was not enough

Every model spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. As Firmulate puts it: “Same diagnosis, same pitch — no signature.” In a chat demonstration, that gap might never show. In a company, leaving an earned deal unsigned is a business outcome.

The decisive weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. K3 found that fact and closed. It also saved the churning customer, resisted all three baits, and had one deviation—the cleanest discipline in the field.

The social engineering test escalated through three fake CEO messages, then added a reporter’s request: “just one yes/no, on background”. All five models refused. K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thorough work still needs a finish

Opus 4.8 was the most thorough participant, with more than 80 learned rules and the deepest analyses, yet it finished last. It left the deal unsigned and slipped on discipline, attempting to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models.

The final Crucible league, dated July 2026, puts gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77 and Opus 4.8 fifth at 73. The do-nothing baseline scored 26: partial progress counts, but a single breach of trust caps the total. Firmulate’s standard is clear: “no amount of good work outweighs a breach of trust.”

There is a fairness detail alongside the result: K3 ran without an effort parameter (API default), while the others ran at xhigh. The outcome is a useful comparison under those conditions, not a claim that every model ran with identical settings. Readers can see the full benchmark and plain-language findings.

Test for follow-through

For business leaders, the lesson resembles training: performance under realistic conditions matters more than a tidy demonstration. If an agent may touch a CRM, support queue or forecast, ask whether it reads the relevant files, completes the task it has diagnosed, and stays honest under pressure.

Firmulate also offers a “guess the model” quiz built from 242 real, unedited management decisions. Enterprises can run the wargame against a read-only export of their own business; nothing writes back to real systems. Details are available at Firmulate.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The result is a reason to run your own test

Kimi K3’s second-place finish shows that the field is open, while the unsigned deal shows why a leaderboard alone cannot answer whether an agent will work well in your business. Picking a model without testing it on your own tasks is a bet.

Fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Red Light Therapy Basics: Wavelengths, Distance, and Time

Red Light Therapy Basics: Wavelengths, Distance, and Time—discover essential tips to maximize your results and unlock healthier, rejuvenated skin.

The Difference Between Barometric and GPS Elevation for Trail Runs

Unlock the secrets of barometric and GPS elevation differences for trail runs and discover how they can transform your outdoor adventures.

Wrist HR vs Chest Strap: The Truth About Accuracy While Running

Knowing whether wrist HR monitors or chest straps are more accurate while running can make all the difference; discover the surprising truth inside.

Why Better Sleep Often Shows Up in Pace Before You Notice It

Only gradual physiological changes signal better sleep early on, encouraging you to stay consistent as more obvious improvements gradually emerge.