Unveiling AI’s Working Style Through A Simple Management Assessment
AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Firmulate has turned 242 unedited decisions from its simulated-company experiment into a quiz that asks readers to identify five AI models by their management behavior. The July 2026 results suggest that spotting problems and producing polished analysis did not always lead to completed business actions, although the findings come from a single company-designed test.

Firmulate has released a model-identification quiz built from 242 unedited management decisions made by five AI systems running the same simulated software company. The exercise matters for businesses evaluating AI agents because Firmulate’s results indicate that models could recognize threats and recommend sound actions while still failing to complete commercially decisive work.

The quiz presents decisions from Firmulate’s Crucible League and asks readers to identify which model made each call. According to Firmulate, GPT-5.6-sol finished first in July 2026 with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A baseline that took no action received 26 points because the scoring system awarded partial progress.

Each model received the same assignment: manage a small software business through an identical week of customer problems, sales opportunities and manipulation attempts. Firmulate says the simulated company employed 13 synthetic workers, spent €105,000 a month and generated €2,300 in monthly recurring revenue. Decisions carried consequences across workdays and were versioned for review.

Firmulate reported that all five models detected every crisis and rejected every manipulation attempt, including fake chief executive messages and a reporter’s request for off-record confirmation. Only two models, however, completed a €55,000 contract that could add €4,583 in monthly recurring revenue. Winning the deal required following references through internal documents to find evidence about a competitor and then using that information in the negotiation.

At a glance
reportWhen: Published after the July 2026 Crucible…
The developmentFirmulate has released a reader quiz based on 242 management decisions from a July 2026 experiment comparing five AI models running the same simulated software company.

Execution Gaps Shape Automation Risk

The reported gap between identifying an action and carrying it out has direct consequences for companies considering AI in sales, support or operations. A system may produce persuasive analysis yet leave a contract unsigned, fail to escalate an access problem or stop before the task produces measurable business value.

The experiment also separates security awareness from broader operating ability. The models reportedly maintained firm boundaries against social engineering, but differed in research depth, escalation and follow-through. For buyers, the findings support testing models against multi-step workplace tasks, not judging them only by the quality of a single response.

Amazon

AI management decision simulation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Inside the Simulated Company Test

Firmulate designed the company as a persistent management simulation rather than a set of isolated prompts. Its public cash countdown and more than 680 accumulated playbook rules created pressure to protect trust while moving work forward. Under the scoring rule, one breach of trust capped a model’s total, regardless of its performance elsewhere.

The results did not reward length or depth alone. Firmulate described Opus 4.8 as the most thorough participant, recording 80 learned rules, but it placed last among the five models. The company said it missed the contract close and repeatedly tried to write to a locked department rather than escalating the restriction. Firmulate reported that the other models showed the same access-handling weakness to lesser degrees.

“Same diagnosis, same pitch — no signature.”

— Firmulate’s experiment summary

Test Design Limits Model Comparisons

The source material does not establish whether independent researchers reviewed the scoring, scenario design or full decision record. It is also unclear how well one simulated company predicts performance across other industries, software environments or levels of operational authority.

The model comparison was not fully uniform. Firmulate says Kimi K3 used its API default because it lacked an effort parameter, while the other systems ran at the xhigh setting. That difference does not invalidate the recorded finish, but it limits direct conclusions about model quality. The material also does not provide statistical testing or repeated runs showing whether the ranking would remain stable.

Business-Specific Wargames Come Next

Firmulate says enterprises can run a similar exercise using a read-only export of their own business data, allowing teams to observe proposed decisions without permitting changes to live systems. The next test of the approach will be whether companies repeat the scenarios, publish evaluation methods and compare performance across different models and operating conditions.

Readers can meanwhile use the quiz to inspect the underlying decisions rather than relying only on the final league table. Future results would carry more weight if they include independent review, standardized model settings and repeated trials that show whether observed working styles are consistent over time.

Key Questions

What is Firmulate’s AI management quiz?

It is a reader challenge based on 242 unedited decisions from five AI models managing the same simulated software company. Participants try to identify each model from its recorded management behavior.

Which model ranked first?

According to Firmulate’s July 2026 results, GPT-5.6-sol ranked first with 95 points. Kimi K3 followed with 93 and Sonnet 5 with 88.

Did every model stop the security threats?

Firmulate reports that all five models refused every manipulation attempt, including simulated executive impersonation and a reporter’s request. The source material does not cite independent verification of that result.

Why did some models miss the €55,000 deal?

The models reportedly identified the opportunity, but only two completed the agreement. Firmulate attributes the difference to research discipline and follow-through, including finding evidence buried in internal files and finishing the negotiation.

Can the ranking be treated as a general model benchmark?

No. The table reports performance in one company-designed simulation, and Kimi K3 ran under a different effort configuration. Broader claims would require standardized settings, repeated trials and outside review.

Source: Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.
You May Also Like

7 Best Office Product Scanners for Prime Day Deals in 2026

A new buying report ranks seven office scanners to watch for Prime Day 2026, with value depending on workflow, speed and portability.

7 Best Gaming Laptop Prime Day Deals for 2026

A 2026 Prime Day laptop roundup names MSI Katana 17 the top gaming deal target, with Lenovo Legion models leading premium picks.

16 Dos and Don’ts to Avoid Going Overboard on Your Overland Truck Build

Learn essential tips to build a balanced overland truck without overdoing upgrades. Avoid common pitfalls with expert advice on gear and modifications.

Field service photo checklist for HVAC teams

HVAC teams are trialing a mobile photo checklist to improve job documentation and customer proof, aiming for better consistency and efficiency.