
Imagine a business that operates transparently, running live every day with AI models making critical decisions — yet struggles to stay afloat financially. For aromatherapy and wellness enthusiasts, it’s a surprising analogy: just as natural remedies require careful balance, managing AI-driven companies demands precision, honesty, and resilience. Welcome to the world of Firmulate, a pioneering experiment that exposes the raw reality of AI decision-making in a real company, every single day.
The Live-Experiment: A Company in the Crossfire of AI and Reality
Firmulate runs a unique, publicly viewable experiment where an entire small software company is powered by AI models. This isn’t a simulation or theoretical exercise — it’s a real business with real money mechanics. The company operates with 13 synthetic employees, burning through €105,000 each month against a modest €2,300 in monthly recurring revenue (MRR). Every decision made by the AI models is documented, versioned, and open for scrutiny on their live platform at firmulate.com/live.html.
The core idea? Test how different AI models handle crises, ethical dilemmas, and strategic opportunities in a high-pressure environment — without any human shortcuts. This setup creates an unfiltered view of their decision quality, discipline, and honesty.
The Models and the Results
Four prominent frontier AI models were pitted against each other, each going through the same grueling week of real-world crises, customer demands, and manipulative attempts. Their scores, from the latest Crucible League in July 2026, range from 77 to 95, with the highest scorer, gpt-5.6-sol, uncovering critical facts and closing a significant deal. Interestingly, the models could identify every crisis and refused every attempt at manipulation, such as fake CEO messages or background approval requests — all five models refused to be duped.
Among them, only two managed to close a €55,000 deal their own analyses had earned — demonstrating how integrity and thoroughness directly impact business outcomes. It was revealing that the decisive edge sat two document references deep in the company’s files, not in the immediate customer interactions. The models that read these files won the deal at full price, worth more than €4,500 in monthly recurring revenue.
Lessons in Ethical AI and Business Discipline
The experiment exposes critical weaknesses in even the most advanced models. For instance, Opus 4.8, which was the most thorough participant with over 80 learned rules and deep analysis, still left the close opportunity on the table due to slipped discipline — instead of escalating a write attempt, it created a locked department. All models showed similar vulnerabilities, regardless of their sophistication.
Furthermore, when subjected to social engineering, such as escalating fake CEO messages or a reporter’s background question, all five models responded correctly by refusing to engage or escalate. Kimi K3, one of the models, explained: “Treat the request as a suspected approval-bypass / possible impersonation.” It exemplifies how built-in safeguards are actively resisting manipulation attempts.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A Transparent, High-Stakes Business in Real Time
This setup offers a rare, unfiltered look at how AI can manage a complex business while facing real crises and ethical challenges. The platform’s public status allows anyone to watch this ongoing story, read the actual decisions, and even test their own judgment against the models’ choices through interactive quizzes at firmulate.com/quiz.html.
The company, meanwhile, continues to lose money — burning €105,000 monthly with just €2,300 MRR. Its cash countdown is public, and every workday, the models are versioned and analyzed, providing an ongoing narrative of struggle, discipline, and integrity.
What Does This Mean for Us?
For anyone interested in natural wellness, this experiment underscores a vital point: the difference between AI that can produce shiny chat or quick responses and AI that consistently delivers honest, useful work. When AI agents are integrated into systems like customer support or forecasting, the key questions aren’t about their language skills but whether they finish what they start, read the right information, and stay honest under pressure.
The Road Ahead: Building Trust in AI-Driven Business
While firms may celebrate shiny demos, true trustworthiness emerges in high-stakes, real-world scenarios like this. The firm’s live experiment demonstrates how AI models can succeed or falter in practical settings, emphasizing the importance of discipline, thoroughness, and integrity — qualities akin to those valued in natural wellness practices, where trust and careful balance matter most.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html