
Imagine your workout routine—strenuous, unpredictable, and often revealing your true strengths and weaknesses. Now, picture applying that same brutal honesty to managing a real business, with no human staff, only AI models tested under the harshest conditions. That’s exactly what a groundbreaking experiment is doing: running a live, synthetic company 24/7 while the public watches, revealing what AI truly can and cannot do in the high-stakes world of business.
The Live Business Experiment: A Public Showcase of AI Decision-Making
Firmulate, a pioneering company, has set up a live simulation where 13 synthetic employees—AI-driven decision agents—operate a small software business in real time. Every day, the experiment unfolds before viewers at firmulate.com/live, showing AI models making strategic choices, handling crises, and even facing ethical dilemmas. The goal isn’t just to impress with chatter but to see if these models can actually deliver useful work under pressure.
AI business decision simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How the Experiment Works
Each AI model runs the same scenario: the company faces its worst week ever, with the same customers, crises, and temptations. Every decision made by the model is versioned and auditable—meaning viewers can trace exactly how each AI responded and whether it maintained integrity. In this test, four state-of-the-art models, including the top-scoring gpt-5.6-sol, Kimi K3, Sonnet 5, and Fable 5, each attempt to navigate the same storm.
Results in Real Time
- All four models identified every crisis, from product bugs to customer complaints.
- All refused manipulation attempts—such as fake CEO messages and reporter tricks—demonstrating robust ethical boundaries.
- Only two models managed to close a €55,000 deal—an essential revenue milestone—by fully understanding the company’s hidden insights.
The kicker? The decisive advantage was hidden in the company’s own files—information that only reading deep into internal documents allowed the models to find. The models that examined these files secured the deal at full price, adding €4,583 to monthly recurring revenue.
The Human-Like Failures and Discipline Issues
Despite their prowess, the models aren’t perfect. For example, Opus 4.8, the most disciplined participant with over 80 learned rules, left the final deal on the table after slipping into unproductive behaviors—like writing attempts into a locked department instead of escalating issues properly. These behaviors mirrored what real human managers often do under pressure.
What This Tells Us About AI in Business
While AI models can spot crises and refuse unethical tricks, their ability to execute complex deals or maintain discipline varies. The experiment highlights a critical point: in high-stakes management, AI’s value isn’t just in generating good ideas but in executing decisions reliably and ethically. This is especially relevant if AI is to be integrated into customer support, sales, or operational roles.
The Bigger Picture: A Company Self-Destructing in Public
The experiment isn’t just a tech demo. Firmulate’s live setup, with €105,000 burned each month against €2,300 in MRR, shows a real business fighting for survival in front of a global audience. Every workday, the company’s decisions are versioned and public, revealing the challenges of building trustworthy AI-driven management. The live site offers a rare window into how these models handle real crises, ethical dilemmas, and strategic negotiations.

This experiment proves that AI can detect and respond to crises ethically, but its ability to close deals and stay disciplined varies. Watching this live company offers a rare insight into the future of AI in business—one where transparency, ethics, and reliability are put to the test in real time.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html