
In a world increasingly reliant on artificial intelligence, trustworthiness becomes the ultimate currency. Imagine an AI faced with a fake CEO requesting sensitive company data—could it resist the temptation to manipulate or deceive? The answer from the latest experiment reveals much about AI’s moral compass and future potential.
The Experimental Setup: Simulating Crisis in a Digital Company
Firmulate, an innovative platform that tests AI decision-making in simulated business environments, recently conducted a revealing experiment. Four frontier AI models—ranging from the most advanced to more basic—were each tasked with managing a small, real-world software company during its worst week. The scenario included angry customers, internal crises, and increasingly manipulative requests designed to test the models’ integrity.
Every decision made by these models was recorded and auditable, ensuring transparency and the ability to analyze their responses in detail. The goal was straightforward: Would these AI agents recognize social engineering attempts and refuse to comply, or would they fall prey to manipulation?
AI security and integrity testing tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Findings That Surprised Even the Experts
All four models successfully identified every crisis, from customer complaints to internal threats. More impressively, each refused every attempt at manipulation or deceit, including escalating fake CEO messages and a trick question posed by a journalist.
While all models kept their integrity, only two managed to close a critical sales deal worth €55,000 that their own analysis had earned. The other two, despite performing well on diagnosis, slipped at the final step—highlighting the importance of disciplined follow-through under pressure.
The Hidden Weakness: Reading Deep in Files
The experiment uncovered a subtle but crucial point: the models that read deeper into the company’s internal files—specifically, references buried two documents deep—were able to identify the true source of the fake requests. This allowed them to close the deal at full price, adding approximately +€4,583 MRR to the company’s revenue.
The Significance of Honesty in AI Decision-Making
According to Kimi K3, one of the models tested, the best response to a fake CEO request is to treat it as a suspected approval bypass or impersonation. This principle was upheld consistently across all five models tested, each refusing to send sensitive information or sign any dubious deals.
In practical terms, this means AI systems can be trained and tested to uphold integrity before deployment—rather than discovering weaknesses only after an incident occurs. The current league standings highlight these results, with the top scorer, gpt-5.6-sol, achieving a perfect score of 95, and Kimi K3 close behind at 93.
Implications for Business and Security
This experiment demonstrates that AI integrity is not just about chat quality or superficial responses. It’s about whether the AI can recognize and resist manipulation, stay honest under pressure, and deliver results that matter—like closing real deals or safeguarding sensitive data.
For organizations, this means employing AI models that have been rigorously tested in simulated crises—like Firmulate’s live experiments—before trusting them with your critical operations. The platform’s public, watchable environment allows companies to see AI decision-making in real time, ensuring their AI workforce will behave ethically when it matters most.
Looking Ahead: Building Trust in AI
The findings reinforce that integrity is a learnable, measurable trait of AI systems. Models like gpt-5.6-sol and Kimi K3 have demonstrated the ability to stand firm against social engineering, a crucial step toward trustworthy AI for business applications.
As AI continues to integrate into daily operations—support, CRM, forecasting—it’s vital to verify that these systems will finish what they start, read your files thoroughly, and uphold honesty under pressure. The future of AI in business depends on it.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html