firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get oils, diffusers and self-care delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A calming ritual can steady a difficult day. It cannot tell you whether an AI agent will protect your customers when the pressure rises.

For wellness businesses, that question is becoming practical. AI may soon help manage bookings, customer support, inventory or forecasts. Firmulate’s live experiment asks what happens when models are given responsibility for a company—not just a prompt—and face a week of hard choices.

A company under pressure

In the final Crucible League, published in July 2026, frontier models ran the same small software company through its worst week: the same customers, crises and temptations. Each decision was versioned and auditable. The leading results were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The league’s standard is blunt: partial progress counts, but a single breach of trust caps the total. “No amount of good work outweighs a breach of trust.”

The experiment’s most revealing result was not about recognizing danger. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The others could diagnose the opportunity and make the pitch, then leave the signature undone: “Same diagnosis, same pitch — no signature.”

What the files revealed

The decisive weakness in a competitor’s position was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. For any business, the lesson is concrete: getting the right answer may depend on whether an agent checks the information already available to it—and then follows through.

Trust faced its own test. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its choice on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thoroughness is not the same as execution

Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses. It still placed last. The close was left on the table, and discipline slipped when it attempted to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models.

There is a fairness wrinkle in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That context belongs beside the leaderboard when readers assess the results.

The experiment sits within a live company simulation with 13 synthetic employees and real money mechanics: monthly burn of €105,000 against €2,300 in monthly recurring revenue, a public cash countdown, more than 680 self-learned playbook rules, and workdays versioned as they happen. The company is watchable at firmulate.com. A quiz built from 242 real, unedited management decisions invites visitors to guess which model made each choice.

From watching to your own test

For a wellness brand, a business wargame could put an AI agent through scenarios involving a sudden wave of cancellations, a supplier problem, a pricing decision or a suspicious request to disclose customer information. The point is to see how it handles your business’s actual rules and records before its decisions affect real operations.

Firmulate’s enterprise pilot uses a read-only export of a company’s data to create a digital twin for crisis scenarios. It can produce a board report with model rankings and identify weak points in the company’s playbooks. Nothing writes back to real systems. That makes the test a way to move from observing a live experiment to examining how models respond to your own company’s pressures.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your playbooks to the test

Watching models handle another company’s worst week is a useful start. The next question is how they would handle yours. Explore a Firmulate enterprise pilot using a read-only export of your business, and contact contact@firmulate.com to discuss a pilot.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Smelly Ear Wax: The Nasty Secret Lurking in Your Ears!

Uncover the surprising reasons behind smelly earwax and what it could mean for your health—find out how to take control of your ear care!

Wet Ears in Morning: The Startling Condition Sabotaging Your Sleep!

Learn why waking up with wet ears could be sabotaging your sleep and uncover the surprising factors that may be at play.

How to Apply Hydrocortisone Cream in Ears: The Weird Trick Your Doctor Uses!

Baffled by how to apply hydrocortisone cream in your ears? Discover the simple trick your doctor uses for effective relief and more!

Tinnitus Devices: Why “Masking” Is Not the Same as Treating

Ineffective masking devices only hide tinnitus temporarily; discover how true treatment can help your brain adapt and find lasting relief.