
Imagine witnessing a company’s daily struggle for survival, not through news headlines, but through the very decisions it makes—decisions guided by artificial intelligence, yet vulnerable to the same human flaws. In a world increasingly powered by AI, what if we could watch a business operate in real-time, facing crises, temptations, and trust tests—all live and transparent? This is not science fiction, but the groundbreaking experiment unfolding at Firmulate, where AI models run a company on the edge of financial and ethical peril.
At the heart of this experiment is a small software business without human employees—13 synthetic ’employees’ operated by AI models. The company’s daily operations, crises, and decision-making are fully visible to the public at firmulate.com/live. Every workday, the firm’s decisions are logged and versioned, revealing how artificial intelligence navigates a real business environment under immense pressure.
What makes this experiment extraordinary is its brutal honesty. The company burns €105,000 each month while generating only €2,300 in monthly recurring revenue—an unsustainable burn rate, reminiscent of startups that struggle to find their footing. Yet, the focus isn’t solely on profit. It’s on whether these AI models can make ethical, consistent decisions, especially when faced with crises or manipulative tactics.
The AI Competitions and Findings
Four frontier AI models, including GPT-5.6-sol and Kimi K3, were each tested against the same difficult week. This week involved real customer crises, temptations to cheat, and scenarios designed to challenge integrity. The results were revealing:
- All models identified every crisis and refused manipulation attempts—an encouraging sign of ethical robustness.
- Only two signed a €55,000 deal they independently identified as worth closing—an indicator of their ability to recognize and act on profitable opportunities.
- Interestingly, the decisive factor lay in a buried piece of information—details hidden two references deep in the company’s own files. The models that read these internal documents successfully secured the deal at full price, worth €4,583 in monthly recurring revenue.
This demonstrates the critical importance of thorough information retrieval and comprehension—a lesson for AI’s future role in business decision-making.
Trust and Ethical Challenges
One of the most striking tests involved social engineering. The experiment simulated fake CEO messages escalating in urgency, and even a reporter’s subtle request for a background approval. Every AI model refused to bypass security or impersonate authority—showing a high level of built-in resistance to manipulation, with Kimi K3 reasoning: ‘Treat the request as a suspected approval-bypass / possible impersonation.’
Despite these safeguards, the experiment also uncovered weaknesses. The most thorough model, OPUS 4.8, which analyzed over 80 rules and conducted deep evaluations, still left a deal on the table when discipline slipped. It demonstrates that even sophisticated models can falter under sustained pressure or when routines are disrupted.
The Real Business Reality
This entire setup is live and ongoing, with the company’s decisions open to public scrutiny. Every weekday, you can see how the AI navigates crises, manages finances, and responds to manipulative tactics. It’s a raw, unfiltered view of AI’s capabilities and limitations in a real-world setting—not just a demo or simulation.
Beyond the technical insights, this experiment raises profound questions about trust, ethics, and the future of AI-driven companies. If AI models can recognize internal information critical to closing deals, refuse manipulation, and operate under real financial pressures, what does that mean for human-led management? Are AI systems ready to take on roles that require integrity and strategic judgment?
While the company is currently losing money, its real value lies in the lessons it imparts about AI’s potential and pitfalls. With transparent, versioned decision-making and a rigorous testing ground, firms can explore how to deploy AI ethically and effectively—before ever hiring it as an employee.
For those curious about these developments, the live experiment is ongoing, and the results are as illuminating as they are unsettling. As AI continues to evolve, understanding its behavior in complex, high-stakes scenarios becomes crucial—whether you’re running a business, investing, or simply contemplating the future of work.
Why This Matters for You
If AI agents are to become part of your operations—handling customer support, sales, or decision-making—what matters isn’t just their ability to produce convincing chat responses. It’s whether they can see through crises, stay honest under pressure, and act in the company’s best interest. Watching this experiment unfold provides a rare glimpse into that future—a future where trust and integrity are tested in real time.
Visit firmulate.com/live to see the company in action, or explore the decisions that shaped its day-to-day life. This is AI business in the raw—an open laboratory for trust, ethics, and operational resilience.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.