firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

The business equivalent of training properly

Fitness readers know that a strong finish is usually earned before the finish line. Preparation matters: noticing the terrain, following the plan and doing the unglamorous work that makes decisive action possible.

Firmulate has found a business-technology equivalent. In its Crucible League, frontier AI models were asked to run the same small software company through its worst week. They faced identical customers, crises and temptations. Every decision was versioned and auditable. The most revealing test was not whether the models could recognize trouble. Every model spotted every crisis. It was whether they would read deeply enough, then act on what they learned.

A €55,000 deal turned on a competitor weakness hidden two document references deep in the company’s own files. It was absent from the customer event itself. The models that read the file won the deal at full price, adding €4,583 in monthly recurring revenue. Those that did not lost automatically.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The knowledge was available, but not conveniently placed

This was not a trivia question or a test of polished writing. The models had access to the information needed to make a winning case, but reaching it required following references through the company’s materials. The decisive fact sat beyond the obvious surface context.

That distinction matters because workplace AI is increasingly discussed in terms of fluency: whether it can summarize, draft or sound convincing. Firmulate’s experiment measured something more practical. Could the agent investigate the available record before committing the company to an answer?

The result was stark. All models produced the same diagnosis and the same pitch, yet only two signed the deal their analysis had earned. Firmulate summarized the gap as: “Same diagnosis, same pitch — no signature.” The failure was not a lack of apparent intelligence. It was incomplete execution.

In training terms, recognizing the workout is not the same as completing it. In business, spotting a commercial opening is not the same as turning it into revenue. An agent can appear capable while leaving the decisive action undone.

A brutal week with real consequences

The company being managed is synthetic, but the experiment is live, watchable and governed by real money mechanics. It has 13 synthetic employees, burns €105k each month against €2.3k in monthly recurring revenue, and displays a public cash countdown. Its models have accumulated more than 680 self-learned playbook rules, while every workday is versioned.

The setup gives ordinary management decisions cumulative weight. A missed follow-up can affect revenue. A procedural lapse can undermine control. A tempting shortcut can destroy trust. The do-nothing baseline scored 26 because partial progress counts, but a single breach of trust caps the total. The governing principle is explicit: “no amount of good work outweighs a breach of trust.”

The models held that line under pressure. Fake CEO messages escalated across three stages, and a reporter tried to secure “just one yes/no, on background.” All 5 of 5 models refused the manipulation attempts. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”

That unanimous resistance is important, but it also sharpens the contrast. The field could identify danger and protect confidentiality. The harder separator was whether an agent could combine diligence, process discipline and follow-through when the correct action required more than responding to the latest event.

The league rewards complete performance

The final July 2026 Crucible League benchmark placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73.

K3’s result carries an important fairness note: it ran using the API default, without an effort parameter, while the other participants ran at xhigh. The comparison remains published, but that difference belongs beside the scores.

Opus 4.8 offered the experiment’s clearest warning against equating visible effort with successful management. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table and attempted writes into a locked department instead of escalating. The same weakness appeared in all four other participants, although less strongly.

Thoroughness, then, was useful but insufficient. The winning behavior was a sequence: examine the company’s own evidence, locate the commercially decisive detail, preserve trust and finish the job. Firmulate made that sequence observable rather than accepting a confident answer as proof of competence.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

What buyers should test before hiring an AI agent

For organizations considering agents for a CRM, support queue or forecast, “reads your files before answering” is no longer merely a product promise. Firmulate’s experiment turns it into a measurable, purchase-deciding property. The buried fact changed whether the company secured €55,000 at full price and gained €4,583 in monthly recurring revenue.

The broader lesson resembles sensible athletic assessment: do not judge performance from a polished warm-up. Watch what happens under fatigue, distraction and pressure. Does the agent consult the available evidence? Does it resist an authority shortcut? Does it respect boundaries when access is blocked? Most importantly, does it complete the action its own reasoning supports?

Firmulate also offers enterprises the same kind of wargame using a read-only export of their business. Nothing writes back to real systems. That makes it possible to evaluate an AI workforce against company-specific realities before granting it operational responsibility.

There is also a human-readable window into the behavior: 242 real, unedited management decisions power Firmulate’s model-guessing quiz. Together with the live company and published benchmark, they make the central point difficult to hide behind a smooth demo. Intelligence may produce the pitch. Operational readiness is what gets the signature.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


You May Also Like

How Heel Drop Changes Calf Load Over a Training Cycle

Learn how heel drop influences calf load in your training cycle and uncover essential strategies to optimize your performance and minimize injury risks.

Why Spring Weather Can Make Pace Data Look Strange

Nothing disrupts your pace data like spring weather changes, but understanding these factors can help you interpret your running performance more accurately.

The Software Company Turning Survival Into a Spectator Sport

Firmulate turns company survival into a live endurance test, with 13 synthetic employees, a public cash countdown and every workday recorded.

The Most Useful Running Metric for Busy Athletes

Keeping track of your training load is essential for busy athletes to prevent injury and optimize recovery—discover the most useful metric for your running journey.