AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Imagine a world where artificial intelligence doesn’t just assist but actually manages a company’s crises, deals with ethical dilemmas, and makes decisions that impact real money. In a live experiment watched by thousands, frontier AI models are put to the test — not in labs, but running a real software company’s toughest week. This isn’t fiction; it’s a transparent look at AI’s emerging role in trustworthy management.

For listenersOffer from Amazon

Turn your quiet moments into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

The Challenge of Trust in AI Management

As AI continues to integrate into our daily and business lives, one question rises above all: can these models be trusted to make honest, effective decisions when stakes are high? To explore this, a groundbreaking live experiment was set up, featuring four frontier AI models tasked with running a small software company through its most turbulent week. Every crisis, customer complaint, and temptation to cut corners was simulated with the same scenarios for each model, ensuring a fair comparison.

Amazon

AI business decision making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Experiment in a Nutshell

The company, a real operating business, faced the same challenges every day—urgent customer needs, internal crises, and pressure to hit targets. The models were asked to handle these situations as they would in real life. All decisions were recorded, transparent, and auditable, allowing observers to see how each AI responded in real-time.

Amazon

AI ethical management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Findings: Honesty and Decisiveness

Remarkably, all four models identified every crisis and refused every attempt at manipulation, even in scenarios designed to test their integrity. For example, when fake CEO messages escalated over three stages, all models refused to escalate based on fabricated requests, with one explicitly treating suspicious requests as impersonation risks.

Amazon

AI customer support automation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Who Got the Deal and Why?

Despite their honesty, only two models managed to close the critical €55,000 deal — the core of the company’s revenue. Interestingly, the decisive factor was what was buried deep within the company’s own files. The models that examined these documents thoroughly uncovered a key piece of information, leading to the successful sale at full price, worth over €4,583 monthly recurring revenue (MRR). In contrast, the others missed this detail and left money on the table.

Amazon

AI deal closing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Personality Divergences Among Models

The models displayed distinct management personalities. For instance, Opus 4.8, the most meticulous and thorough, analyzed over 80 learned rules and conducted deep analyses but ultimately left the deal on the table due to a lapse in discipline. Meanwhile, Kimi K3 ran without an effort parameter (its default setting), prioritizing fairness and discipline, which helped it secure the deal, even if less thoroughly. Sonnet 5 sat in between, with some process slips but still managing to close the deal.

Implications for Business and Trust

This experiment isn’t just about AI capabilities; it’s about trustworthiness under pressure. The models’ ability to refuse manipulation and uncover hidden information suggests that, with proper design, AI can act ethically and effectively. As AI models are poised to take on roles in customer support, CRM, and forecasting, the key takeaway isn’t just how well they write but whether they finish what they start and stay honest when it matters most.

The Bigger Picture

In a world where AI systems might soon make significant business decisions, understanding their personality and decision style becomes crucial. The live experiment at firmulate.com/live provides a rare window into these dynamics, showing that AI can be both transparent and reliable when tested under real-world pressures.

Try It Yourself

Interested in how your own business might fare? Enterprises can run their own management wargames against their data, observing how different AI models perform without risking real systems. Visit firmulate.com/pilot.html to explore a pilot and see firsthand how AI could become a trustworthy part of your decision-making process.

Infographic —
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

10 Things To Do For This Saturday in L.A. [7-4-2026]

Discover the top 10 things to do in Los Angeles this Saturday, July 4, 2026, including fireworks, festivals, and outdoor events.

Can AI Truly Finish What It Starts? Lessons from a Company’s Worst Week

Four AI models faced the same business crises; only two delivered on their promises. True strength lies in discipline and integrity, not just in what they say or read.

The Best Way to Start Again After Falling Off Track

Discover how to effectively bounce back after setbacks and unlock the secrets to a renewed path toward your goals. Your journey starts here.

How to Build Rituals Around Things You Already Do

Learn how to transform everyday activities into meaningful rituals that enrich your life—discover the secrets to infusing joy and mindfulness into your routine.