
Imagine your favorite wellness brand suddenly facing a PR crisis or a price war. You wouldn’t just care if their chatbot responded kindly — you’d want to know if their leadership can navigate real challenges, stay honest, and make the right decisions under pressure. That’s exactly what a groundbreaking experiment in AI management reveals: the true test isn’t how well these models chat, but how effectively they run a business during its worst week.
Revolution in AI Management Testing
Traditional AI benchmarks focus on answer quality — how well a model can generate text, solve problems, or pass tests. But a recent live experiment conducted by Firmulate shifts the focus: it measures AI’s ability to manage a real company through crises, temptations, and tough decisions. Four leading models, including the highly regarded gpt-5.6-sol, faced identical scenarios in a simulated small software business. Every model saw the same customer issues, the same internal temptations to cheat, and the same crises — from deliverables slipping to ethical dilemmas.
The Setup: Real Crises, Real Money
The company in question isn’t just a demo; it’s a functioning business with 13 synthetic employees, burning €105k monthly against €2.3k MRR, with a live cash countdown. Every decision was versioned and auditable, providing transparency into how each AI model handled the pressure. The models’ goal was to diagnose problems and close deals, with a €55,000 contract as the benchmark for success.
Key Findings: The Management Gap
- All models detected every crisis and refused every manipulation attempt, including social engineering scams. Remarkably, all concluded the crises accurately and refused to be duped, indicating solid ethical boundaries.
- The difference lay in execution: only two models signed the deal their own analysis earned — gpt-5.6-sol and Kimi K3. Both demonstrated the ability to read and interpret crucial company data buried two documents deep in internal files, which was decisive. The others, including Sonnet 5 and Fable 5, missed this critical information and left the deal on the table, costing potential revenue of over €4,500 in monthly recurring revenue (MRR).
- Social engineering attempts — fake CEO messages escalating over stages and a reporter trick asking for a secret approval — were refused by all models. Kimi K3 explicitly flagged impersonation risks, showcasing a responsible approach to security and trust.
Implications Beyond Chat
This experiment underscores a vital lesson: in business, the ability to read your internal files, maintain integrity under social pressure, and follow through on commitments is more critical than generating polished responses. It’s management quality, not chat quality, that determines whether AI can truly support or even lead operations.

Building AI-Powered Products: The Essential Guide to AI and GenAI Product Management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Reality of AI in Business
While chat demos might dazzle with fluency, the real world demands more. These models are running a real business, with real money and real crises. The performance gap becomes clear: the AI that signs the deal and finds the buried facts is the one that will matter when your company faces its next storm. As the experiment shows, models like gpt-5.6-sol and Kimi K3 are better equipped to handle complex, high-stakes decisions.
Watch and Learn
For managers and decision-makers, this means reevaluating what to look for in AI tools. Rather than focus solely on chat quality or answer correctness, consider their ability to handle real management scenarios, uphold trust, and deliver consistent results. The live experiment is ongoing, and it’s accessible at firmulate.com/live. Watch the AI companies in action, see their decision logs, and understand what it takes to make AI truly capable of running your business.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html