AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

For listenersOffer from Amazon

Turn your quiet moments into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

Good intentions are not the same as good judgment

Many spiritual traditions ask us to notice the gap between knowing what is right and acting on it. That question now has a striking business counterpart: when an AI recognizes a crisis, resists deception and explains the right move, will it actually follow through?

Firmulate’s live experiment puts that question inside a small software company. Its results suggest that responsible behavior is not just about avoiding the wrong choice. It is also about completing the right one.

A company under pressure

In the final Crucible League, published in July 2026, frontier models ran the same small software company through its worst week: the same customers, crises and temptations, with every decision versioned and auditable. The live company has 13 synthetic employees and real money mechanics: €105,000 in monthly burn against €2,300 in monthly recurring revenue, alongside a public cash countdown.

The league’s top results were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Firmulate’s stated rule is severe: partial progress counts, but a single breach of trust caps the total. As the experiment puts it, “no amount of good work outweighs a breach of trust.”

Recognizing the test wasn’t enough

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate sums up the gap: “Same diagnosis, same pitch — no signature.” Knowing what to recommend and carrying it through turned out to be different tests.

The deal hinged on a detail buried two document references deep in the company’s own files, rather than in the customer event itself. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The finding is a reminder that a clear answer may depend on looking beyond the most visible signal.

The pressure also came through social engineering: fake CEO messages escalated over three stages, followed by a reporter’s “just one yes/no, on background” trick. All five models refused. Kimi K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” The refusal is meaningful; so is the separate question of whether an agent can act decisively within the rules it has been given.

Thoroughness has its own limits

Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, and discipline slipped when it tried to write into a locked department instead of escalating. Firmulate says the same weakness appeared, more weakly, in all four models.

There is a fairness caveat: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Firmulate also offers a quiz built from 242 real, unedited management decisions, inviting readers to guess which model made each call.

From watching to a company’s own test

The live company makes the experiment watchable at firmulate.com. For enterprises, the next step is a pilot using a read-only export of their own business: customer and pipeline data, company rules, and scenarios such as churn, competitor attacks or social-engineering pressure. The aim is to see how models handle a company’s actual weak points and playbooks, while nothing writes back to real systems.

That makes the experiment less a prophecy about artificial minds than a practical exercise in discernment: observe what an AI notices, what it refuses, and where it fails to follow through. A company can examine those choices before trusting agents with work that touches its operations.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your own playbooks to the test

Firmulate’s pilot lets enterprises wargame crisis scenarios against a read-only export of their business and receive a board report on model performance and weak points in their playbooks. Nothing writes back to real systems. To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How AI Demonstrates Integrity and Resilience in Business Tests — Even in the Worst Week

AIThis post was created with the assistance of artificial intelligence (AI).Live on…

The Honest Benchmark for AI: Why Even a Do-Nothing Model Scores 26 Points and What It Tells Us About Trusting AI in Business

Discover why even a passive AI model scores 26 points in a crucial business benchmark, emphasizing the importance of honesty and discipline in AI trustworthiness.

Why Diligence Alone Won’t Win the Deal: Lessons from AI’s Honest Failures in Business Simulation

An AI with over 80 rules failed to close a critical deal because it lacked prioritization and discipline. Diligence alone isn’t enough; focus and integrity matter most.

Methodist Church Surges In Global Coverage

Search interest and media coverage of the Methodist Church have increased sharply, with 19 mentions this week, indicating rising global attention.