
Imagine an AI that does nothing—yet still earns points, hits scores, and passes complex tests designed to mimic real-world business crises. Sounds counterintuitive? It’s a deliberate approach to measuring AI trustworthiness and reliability, and it reveals vital insights about what we should really expect from AI systems, especially those managing sensitive or critical tasks.
Get health and wellness essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
In the quest to develop AI capable of managing real business operations, researchers have devised a transparent and rigorous benchmarking process. The latest experiments from Firmulate, a leader in AI management testing, place a spotlight on something surprising: even a ‘do-nothing’ AI baseline scores 26 points out of a possible 100. This score is not zero—why? Because partial progress and inherent safeguards are factored into the scoring system.
The experiment involved running multiple frontier AI models through the same challenging business week—complete with customer crises, internal temptations, and manipulative tactics. These models had to navigate the complexities of decision-making, trust, and honesty. Interestingly, all models succeeded in identifying every crisis and refused all manipulative attempts, but only two managed to close a key deal and sign a €55,000 contract, based on their own analysis and decision-making.
So, where did the difference lie? The decisive factor was the models’ ability to read and interpret company documents. The models that examined internal files found critical information buried two document references deep, which enabled them to close the deal at full price—adding more than €4,500 in monthly recurring revenue (MRR). The others, despite understanding the situation, left the opportunity on the table, demonstrating that deep information processing can be the differentiator in trust and performance.
One of the most revealing aspects of this experiment is how the models handled social engineering attempts—fake CEO messages escalating over three stages and a reporter trick. All five models refused to escalate or approve these manipulative requests, with Kimi K3 explicitly treating such requests as potential impersonation or approval-bypass attempts. This consistent refusal underscores the importance of built-in safeguards against trust breaches, especially in real-world applications where manipulation could be costly or dangerous.
The live experiment runs at firmulate.com/live showcase an operational company—13 synthetic employees managing real money mechanics, with a burn rate of €105,000 per month against a revenue of €2,300. The system employs over 680 self-learned rules, versioned daily, illustrating how AI models can be trained, tested, and trusted in a controlled environment before being deployed in actual business settings.
Among the models tested, Opus 4.8—considered the most thorough with over 80 learned rules—performed the worst. It left the close opportunity unexploited, slipping into departmental silos instead of escalating issues properly. This indicates that even deep, rule-based AI can falter if its discipline erodes under pressure, highlighting that more rules and analyses don’t necessarily guarantee better outcomes.
This entire benchmarking process is designed to be transparent and repeatable. Participants can run their own scenarios against read-only exports of their business models, ensuring that AI performance is measurable and comparable without risking real systems. The goal is to prepare AI for the nuanced, trust-dependent decisions necessary in health and wellness sectors, where dishonest practices or manipulative tactics could have serious consequences.
What should health and wellness operators take away from this? The focus isn’t just on how well an AI writes or converses; it’s on whether it can finish what it starts, interpret critical documents, stay honest under pressure, and deliver real, measurable work—especially when human oversight is limited. The Firmulate benchmarks champion a new standard: trustworthiness, transparency, and discipline are the true measures of a successful AI workforce.

For health and wellness providers, the key lesson from this AI benchmark is clear: the value of trust and reliability in AI systems goes beyond surface-level capabilities. An AI that can read deeply, uphold honesty, and complete tasks accurately is essential—especially when lives and well-being are at stake. Firmulate’s transparent testing shows that even a do-nothing baseline scores 26, underscoring that partial progress and safeguards matter most in deploying responsible AI.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI trustworthiness benchmarking tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
business AI decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
