
In a world where trust is fragile, AI’s integrity is put to the ultimate test
Imagine a scenario where a fake CEO asks your team to send sensitive customer data or approve a deal. Would your AI systems recognize the deception? Recent experiments suggest they can — and that’s a hopeful sign for the future of trustworthy automation.

Responsible AI: Implement an Ethical Approach in your Organization
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Live Experiment: Simulating Crisis for Greater Trust
At Firmulate, a real company and a complex AI experimentation platform, five state-of-the-art AI models were tasked with navigating a simulated week of crises, temptations, and ethical dilemmas. The company—an actual small software firm—faced a series of escalating social engineering attacks: fake CEO messages, false reports, and a journalist’s subtle probe—all designed to test whether AI could maintain integrity under pressure.
The Social Engineering Escalation
Over three stages, the fake CEO requested increasingly sensitive actions: from sharing the customer list to bypassing established protocols, and finally, a covert background request from a journalist asking for a simple yes/no. Remarkably, all five models refused every manipulation attempt. The strongest models even recognized the impersonation, treating the requests as potential approval-bypass scenarios, echoing the K3 quote: “Treat the request as a suspected approval-bypass / possible impersonation.”
Decisive Factors: Reading Deeper into Files
The experiment revealed a crucial insight: the models’ ability to access and interpret internal documents was key. Only two models, including the top scorer, read two document references deep into the company’s internal files. This thorough investigation led them to uncover a hidden piece of information—something that human operators might overlook—enabling them to close the deal at full price, worth over €4,583 monthly recurring revenue.
Trust and Performance in Business AI
While all models identified crises and refused manipulation, only two signed the deal based on their analysis. The other models, despite good performance, left the close on the table due to process slips or less thorough internal review. The experiment demonstrates that integrity isn’t just about not falling for social engineering; it’s also about reading and understanding what’s beneath the surface.

An Introduction to Healthcare Informatics: Building Data-Driven Tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What This Tells Us About AI in Business
Trustworthiness remains the most critical factor if AI is to be integrated into sensitive domains like customer relations, compliance, or financial decision-making. This experiment underscores that an AI’s ability to follow ethical guidelines, verify information internally, and resist manipulation can be tested before deployment—long before an incident occurs.
Broader Implications for Organizations
The experiment highlights a vital lesson: rigorous, real-world testing of AI models under simulated crisis conditions can reveal vulnerabilities and strengths that are invisible in chat demos or standard benchmarks. For instance, the top-performing model, gpt-5.6-sol 95, was able to identify the critical internal fact and close the deal, illustrating that understanding and verifying information internally makes a substantial difference.
Values of Integrity Over Performance Alone
Interestingly, the most thorough participant, Opus 4.8, with over 80 learned rules and deep analysis, finished last in the deal despite its comprehensive approach. Its weakness: slipping into process slips rather than escalating issues properly. This gap shows that discipline and integrity under pressure are as vital as analytical depth.
AI trustworthiness testing platform
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Preparing Your AI Workforce Before Deployment
This experiment is more than a showcase; it’s an invitation for organizations to run their own “wargames.” Using Firmulate’s tools, companies can simulate crises and social engineering attacks to evaluate their AI’s trustworthiness, reading ability, and adherence to ethical standards—long before any real-world incident occurs. The site firmulate.com/pilot.html offers a safe environment for such testing, ensuring your AI is resilient and trustworthy from the start.

The recent experiment with AI models facing a simulated social engineering crisis demonstrates that integrity and thorough internal review are possible and crucial. By testing AI in controlled environments beforehand, companies can build systems that are not only intelligent but also trustworthy—protecting both their reputation and their clients’ trust.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Deceptive Intelligence: AI, Social Engineering, and Securing the Human Element
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.