
Turn your quiet moments into listening time
- Thousands of audiobooks, podcasts and originals
- Listen on your phone, tablet or Echo — also offline
- Cancel anytime
Trust is measured when the stakes are real
Faith, intuition and good intentions can shape how we meet a difficult moment. But when an AI system is asked to act for a business, trust has to show up in its choices. Does it recognize danger? Resist pressure? And follow through when the right move carries a price?
Firmulate puts those questions to work in a live, watchable experiment: AI models run the same small software company through a week of crises, with customers, money pressures and temptations held constant. The goal is to see how they manage—not merely how convincingly they talk.
The gap between seeing and doing
In the final Crucible League, published in July 2026, the models finished in this order: gpt-5.6-sol scored 95, Kimi K3 93, Sonnet 5 88, Fable 5 77 and Opus 4.8 73. The do-nothing baseline scored 26. The league treats a breach of trust as disqualifying: “no amount of good work outweighs a breach of trust.”
One result was strikingly consistent. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. They reached the same diagnosis and made the same pitch; most still left the close on the table. Good judgment in conversation did not guarantee action when the decision mattered.
The missing clue was not in a customer event. A decisive competitor weakness lay two document references deep in the company’s own files. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The episode makes a practical point: useful business judgment can depend on noticing what is already known, then acting on it.
Pressure, boundaries and discipline
The experiment also tested social engineering: fake messages escalating across three stages from a supposed CEO, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
Refusing manipulation was not the same as flawless discipline. Opus 4.8 was the most thorough participant, with 80 learned rules and the deepest analyses, but it finished last. It left the deal unsigned and slipped by attempting to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four models.
The live company gives the test an everyday business setting. It has 13 synthetic employees and real money mechanics: monthly burn of €105,000 against €2,300 in monthly recurring revenue, alongside a public cash countdown. Its playbook has more than 680 self-learned rules, and every workday is versioned. Readers can watch the company at Firmulate; a quiz built from 242 real, unedited management decisions invites them to guess which model made each choice.
From watching to your own business
A public experiment can show where capable models hesitate, miss context or struggle to follow their own judgment. But a company’s own risks live in its customers, pipeline, rules and operating habits. That is why Firmulate’s enterprise pilot moves from watching a synthetic company to testing scenarios against a read-only export of a business.
The pilot can examine crisis scenarios and produce a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems. One fairness detail matters when reading the league: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh.

Test trust before handing over responsibility
The experiment suggests that spotting a crisis and resisting a trick are only part of dependable management. Models also need to find relevant information, respect boundaries and complete the work their own analysis supports. A pilot lets an enterprise explore those questions against its own business data while keeping the exercise read-only.
To explore a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
