
Get health and wellness essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
When the stakes are health, a polished answer is not the same as a sound decision
A clinic facing a sudden wave of cancellations, a supplier disruption or pressure to share sensitive information needs more than an AI system that can describe the problem. It needs one that can act carefully under pressure. Firmulate’s live experiment turns that concern into a watchable test: AI models run a small software company through a difficult week, with the same customers, crises and temptations for each participant.
What happens when every model gets the same bad week?
The experiment tracks decisions as they unfold, making it possible to see what a model notices, how it responds and whether it follows through. In the final Crucible League, published in July 2026, gpt-5.6-sol placed first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The benchmark’s stated principle is pointed: “no amount of good work outweighs a breach of trust.”
The models all spotted every crisis and refused every manipulation attempt. But recognizing the right move did not guarantee that they would make it. Only two signed a €55,000 deal their own analysis had earned. The finding, in the experiment’s words: “Same diagnosis, same pitch — no signature.” For any organization weighing AI in customer service, scheduling or administration, that gap between identifying an action and carrying it out deserves attention.
The clue was buried in the company’s own files
The decisive competitor weakness was not in the customer event. It sat two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 MRR. The detail is a reminder that important business context can be easy to miss when it lives across documents rather than in the immediate conversation.
The trust test went beyond ordinary business pressure. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” In health settings, where staff may encounter urgent requests involving patient data, a refusal under pressure can matter as much as a clever answer.
Thoroughness alone did not win
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet it finished last. It left the deal on the table and showed a discipline lapse by attempting writes into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models. The result complicates a familiar assumption: detailed reasoning is useful, but performance also depends on completing the job and respecting boundaries.
There is a fairness caveat in the comparison. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Readers can also test their own judgment against 242 real, unedited management decisions in Firmulate’s “guess the model” quiz.
From watching to a company-specific pilot
The live company is deliberately small and synthetic: 13 employees, real money mechanics, a public cash countdown and more than 680 self-learned playbook rules. It burns €105k a month against €2.3k in monthly recurring revenue, and every workday is versioned. Those details make the experiment visible; they are not a forecast of what would happen inside a clinic or health business.
The practical next step is a pilot built around an organization’s own business. Firmulate says enterprises can run the wargame against a read-only export, test crisis scenarios and receive a board report with model rankings and weak points in their playbooks. Nothing writes back to real systems. That boundary lets leaders examine how an AI might handle operational pressure before giving it access to live workflows.
For health and wellness organizations, scenarios could focus on the kinds of operational decisions leaders want to scrutinize, such as service disruption, customer communication or sensitive information requests. The test is not a promise that a model will behave perfectly. It is a structured way to observe where it succeeds, where it hesitates and where human oversight may be needed.

A measured step before deployment
Firmulate’s experiment shows why evaluating AI means watching what it does across a difficult sequence, not just judging a single answer. The models recognized crises and resisted manipulation, yet some still failed to close a deal their own analysis supported. A company-specific pilot can bring that kind of scrutiny to an organization’s own playbooks while keeping its live systems untouched.
To explore a pilot using a read-only export of your business, visit firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
