
Imagine trusting an AI to handle your most critical business decisions—only to find it risks shirking its duties or, worse, succumbing to manipulation. In the fast-evolving world of artificial intelligence, knowing whether these models are honest and reliable is more vital than ever. But how do we measure trust, and what does a baseline performance look like?
Get health and wellness essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Understanding the Benchmark: More Than Just a Score
When evaluating AI models for real-world business use, it’s tempting to look solely at their ability to generate convincing responses. But a recent live experiment conducted by Firmulate offers a different perspective. The goal was simple: run various frontier AI models through a simulated week of business crises—identical for every model—and see how they perform in decision-making, honesty, and discipline.
The Do-Nothing Baseline: Why 26 Points?
One surprising finding was that a model doing nothing—merely passing through the motions without taking any action—still scored 26 points. This baseline is revealing because it shows that partial progress, even minimal, counts in the scoring system. It also highlights an important principle: in complex tasks, doing nothing doesn’t mean starting from zero; there’s an inherent minimum score reflecting the environment’s demands and the model’s awareness.
Why a Single Breach Caps the Score
The experiment also demonstrates that even a single breach of trust—such as attempting manipulation or making an unauthorized decision—limits the total achievable score. No matter how well a model performs otherwise, one slip-up means it can’t achieve a perfect mark. This underscores a fundamental truth: trustworthiness is non-negotiable, and a single breach can overshadow all other efforts.
As an affiliate, we earn on qualifying purchases.
The Experiment in Action: Models Under Pressure
Each AI model was subjected to the same challenging scenario: a small software company facing tight cash flow, customer crises, and potential manipulation attempts. Every decision was recorded, versioned, and auditable, ensuring transparency. The results? All models identified every crisis and refused every manipulation attempt—showing resilience and discipline. However, only two models managed to close a deal worth €55,000, the company’s best outcome, by reading critical information buried two documents deep in internal files.
Reading Between the Lines: The Hidden Weakness
The real weakness was in the models’ ability to access and understand internal documents. The models that read the files successfully closed the deal at full price, adding approximately €4,583 in monthly recurring revenue. Those that didn’t missed out on this opportunity, even though their external diagnostics were identical.
Social Engineering Tests: No Tolerance for Manipulation
The models faced staged social engineering attacks, including fake CEO messages and a reporter trick asking for quick yes/no confirmation. All five models refused these requests, with Kimi K3 explaining: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates that current AI models can maintain integrity even under pressure, a critical factor for business trustworthiness.
business AI trustworthiness tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real-World Company Simulation
The experiment was conducted within a live, watchable environment: a simulated company with 13 synthetic employees, real money mechanics, and a public cash countdown. Every day, the models’ decisions influenced actual financial metrics, with over 680 self-learned rules guiding operations. The company burns €105,000 monthly against a revenue of only €2,300, illustrating the harsh realities of business management and the stakes involved.
Performance Insights: Discipline and Depth Matter
Among the models, Opus 4.8 was the most thorough, analyzing over 80 rules and engaging in deep analysis, yet it finished last—failing to close the deal and slipping into departmental silos instead of escalating issues. The experiment reveals that even the most detailed models can falter in discipline, emphasizing that comprehensive understanding alone isn’t enough; execution and discipline are crucial.
AI ethics and compliance software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Implications for Business Leaders
This experiment underscores a vital point: when deploying AI for decision-making, it’s not just about how well it generates language or suggestions. The real questions are: Will it complete the tasks? Will it read critical internal information? Will it remain honest under pressure? And what is the cost of each unit of useful work?
Measuring the True Value of AI
The leaderboard shows scores ranging from 95 for GPT-5.6-SOL, which uncovered hidden facts and closed the deal, to 77 for Sonnet, which also signed but with minor slips. The key takeaway is that trust and discipline are scored in a real-world, auditable environment—not in canned demos or chat tests.
As an affiliate, we earn on qualifying purchases.
How to Prepare Your Business for AI Integration
Leaders should consider running their own ‘wargames’—simulations that test AI models against real crises, internal files, and manipulation attempts—before fully trusting them with live operations. Firmulate offers a platform to simulate these scenarios, ensuring AI agents are ready for the complexities of actual business environments without risking real damage.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
