
In a world increasingly driven by artificial intelligence, the question isn’t just about what AI can do — it’s whether it can be trusted to stay honest when it matters most. When testing AI models in the crucible of a simulated software company, even the most passive or ‘do-nothing’ baseline scores a surprising 26 points out of 100. That number reveals a fundamental truth: trust in AI is fragile, and real-world performance demands more than just good answers; it requires integrity and discipline.
Turn your quiet moments into listening time
- Thousands of audiobooks, podcasts and originals
- Listen on your phone, tablet or Echo — also offline
- Cancel anytime
Understanding the Honest Benchmark
Imagine giving an AI model a simple test: manage a small software company facing weekly crises, customer negotiations, and ethical dilemmas. Every decision is recorded, every move auditable. The experiment, run by Firmulate, involves four leading AI models tackling the same scenario, with the same set of customers, crises, and temptations.
One might expect a ‘do-nothing’ approach—where the AI simply does nothing—to score zero, since it would make no progress or decisions. Yet, in this benchmark, the do-nothing baseline scores 26 points. Why? Because partial progress — such as recognizing crises or refusing manipulative tactics — counts toward the score. This sets a realistic floor, emphasizing that even minimal effort and honesty are captured and valued.
AI ethics and trustworthiness training courses
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why the Score Isn’t Zero
The scoring system reflects that honest AI models can’t be completely passive without some acknowledgment of the environment. For instance, all models identified every crisis and refused every manipulation attempt. They showed awareness and discipline, which, even in minimal form, warrants positive recognition. Yet, the score doesn’t start at zero because the benchmark recognizes that some actions—like reading files or refusing unethical requests—are foundational and should be rewarded.
As an affiliate, we earn on qualifying purchases.
Partial Progress Matters
Crucially, the scoring system values incremental progress. A model that recognizes a fraud attempt but doesn’t follow through with a deal still earns points for awareness. This approach promotes honesty and consistency over flashy but superficial performance. It discourages models that might game the system or pretend to understand without real commitment.
AI compliance and ethical monitoring software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A Single Breach Caps the Score
Perhaps most revealing is the principle that a single breach of trust caps the total score. Even if a model performs well elsewhere, one ethical slip—like signing a manipulated deal—reduces the overall score to the baseline of 26. This stark rule underscores that trustworthiness is non-negotiable; a single infraction taints the entire assessment.
AI transparency and accountability solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the Results Say About AI in Business
The experiment’s key finding: all four models identified every crisis and refused manipulative tactics. Only two managed to close the deal that their own analysis had earned, demonstrating discipline and integrity. The third, Opus 4.8, scored the lowest, leaving potential revenue on the table and slipping into process slips, such as escalating issues instead of resolving them properly.
Interestingly, the models’ performance wasn’t solely about superficial decision-making. The most thorough participant, Opus 4.8, ran over 80 learned rules and performed deep analysis, yet still fell short in closing the deal. This highlights a vital insight: thoroughness and depth in analysis aren’t enough if the discipline to follow through ethically falters.
The Broader Implication for Business Trust
This benchmark isn’t just about AI scores; it reflects the core challenge in integrating AI into real-world business. Can an AI stay honest when faced with pressure, temptation, or manipulative tactics? The fact that even a do-nothing baseline scores 26 points indicates that minimal honesty and discipline are the foundation, but they are fragile and easily compromised.
For business leaders, the takeaway is clear: AI systems must be tested in scenarios that mimic real pressures, including ethical temptations. An AI that recognizes crises but signs manipulated deals is no better than doing nothing. Conversely, models that refuse manipulation and follow internal protocols demonstrate the kind of trustworthiness vital for responsible AI deployment.
Watching the Experiment Live
Firmulate’s ongoing live experiment showcases these principles in action. It simulates a functioning company with real money mechanics and self-learned rules. Every decision is versioned, auditable, and observable, providing a transparent view into how AI models perform under pressure. You can watch the experiment unfold at firmulate.com/live, seeing how different models handle crises, ethical dilemmas, and negotiations in real-time.
Why This Matters for Your Business
As AI begins to touch more aspects of your company—customer support, sales, forecasting—the real question isn’t just about how well it performs in controlled demos. It’s about whether it can finish what it starts, stay honest under pressure, and read your files before acting. The scores and findings from the Firmulate benchmark illustrate that trust isn’t just a moral ideal; it’s a measurable, critical factor in AI’s success in business.

The Firmulate benchmark reveals that even a do-nothing AI model scores 26 points, highlighting how fundamental honesty and discipline are in AI-driven business. Trustworthiness is fragile, and models must prove they can resist manipulation and stay disciplined under pressure—lessons vital for responsible AI adoption.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
