
In the world of AI, trust is everything—yet even the most minimalist algorithms score above zero. What does this tell us about the integrity of AI assessments and their implications for your business decisions? At the heart of the matter lies a simple truth: even an AI that does nothing earns a baseline score of 26. This isn’t an oversight; it’s a reflection of a rigorous, transparent process that reveals just how much honesty and diligence matter in AI performance.
Turn your quiet moments into listening time
- Thousands of audiobooks, podcasts and originals
- Listen on your phone, tablet or Echo — also offline
- Cancel anytime
Understanding the Baseline: Why 26?
When testing AI models in complex, real-world simulations, the goal isn’t simply to see if they can generate impressive responses. Instead, the focus is on how reliably they manage tasks, make decisions, and maintain integrity under pressure. The experiment conducted by Firmulate placed four advanced AI models—each representing the cutting edge—inside a simulated small software company facing a challenging week filled with crises, manipulative threats, and ethical dilemmas.
Remarkably, even the least active model, dubbed the ‘do-nothing baseline,’ scored 26 points out of a potential 100. This score isn’t a fluke or a starting point; it’s a carefully determined minimum that accounts for partial progress. In other words, even a minimally engaged AI, that perhaps does almost nothing, will earn some points for simply recognizing scenarios and refraining from obvious misconduct. This scoring approach underscores the importance of honesty and cautious decision-making—vital traits for AI systems in business settings.
As an affiliate, we earn on qualifying purchases.
The Role of Trust and Accountability in AI
One striking finding from the experiment is that models which read deeper into a company’s documents—beyond surface-level data—secured better outcomes, including closing high-value deals. For instance, a competitor that examined two document references in the files ultimately won a deal worth over €4,583 MRR. Conversely, models that only relied on surface information or failed to verify internal data missed these opportunities.
Moreover, the models’ behavior under social engineering attempts—like fake CEO messages and media tricks—was telling. All models refused to be manipulated, demonstrating a shared commitment to ethical decision-making. Kimi K3’s reasoning, for example, was to treat suspicious requests as potential impersonation, highlighting the importance of built-in safeguards against deception.
Why Partial Progress Matters
The experiment also shows that partial successes count, but only up to a point. Even if an AI recognizes crises and refuses manipulation, slipping in internal processes—such as leaving deals on the table—reduces its overall score. The most thorough participant, Opus 4.8, with over 80 learned rules and deep analysis, still finished last due to discipline lapses. This delicate balance underscores that in real business, honesty, discipline, and thoroughness are indispensable.
Implications for Business and AI Trustworthiness
For executives and decision-makers, these findings are more than academic; they’re a call to scrutinize AI performance not just by how well it chats, but by how well it acts. Will it follow through on commitments? Will it read and interpret critical internal documents? Will it stay honest even when tempted? These are the questions that truly matter when integrating AI into vital operations.
The live experiment by Firmulate, available at firmulate.com/live, makes this transparent and watchable—an open window into AI’s real capabilities and limitations. The experiment involves running AI models through scenarios that mirror the pressure and temptations of actual business, measuring not just intelligence but integrity and discipline.

The core lesson: even the simplest, do-nothing AI scores 26 points, reflecting the importance of honesty, discipline, and thoroughness. Trustworthy AI isn’t just about answering well; it’s about acting with integrity when it counts most. Firms that understand this will better navigate the complex landscape of AI adoption, ensuring their tools serve their values and goals—rather than just delivering impressive but superficial responses.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
