AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine trusting an AI to handle your most critical business decisions—only to find it risks shirking its duties or, worse, succumbing to manipulation. In the fast-evolving world of artificial intelligence, knowing whether these models are honest and reliable is more vital than ever. But how do we measure trust, and what does a baseline performance look like?

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get health and wellness essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Understanding the Benchmark: More Than Just a Score

When evaluating AI models for real-world business use, it’s tempting to look solely at their ability to generate convincing responses. But a recent live experiment conducted by Firmulate offers a different perspective. The goal was simple: run various frontier AI models through a simulated week of business crises—identical for every model—and see how they perform in decision-making, honesty, and discipline.

The Do-Nothing Baseline: Why 26 Points?

One surprising finding was that a model doing nothing—merely passing through the motions without taking any action—still scored 26 points. This baseline is revealing because it shows that partial progress, even minimal, counts in the scoring system. It also highlights an important principle: in complex tasks, doing nothing doesn’t mean starting from zero; there’s an inherent minimum score reflecting the environment’s demands and the model’s awareness.

Why a Single Breach Caps the Score

The experiment also demonstrates that even a single breach of trust—such as attempting manipulation or making an unauthorized decision—limits the total achievable score. No matter how well a model performs otherwise, one slip-up means it can’t achieve a perfect mark. This underscores a fundamental truth: trustworthiness is non-negotiable, and a single breach can overshadow all other efforts.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Experiment in Action: Models Under Pressure

Each AI model was subjected to the same challenging scenario: a small software company facing tight cash flow, customer crises, and potential manipulation attempts. Every decision was recorded, versioned, and auditable, ensuring transparency. The results? All models identified every crisis and refused every manipulation attempt—showing resilience and discipline. However, only two models managed to close a deal worth €55,000, the company’s best outcome, by reading critical information buried two documents deep in internal files.

Reading Between the Lines: The Hidden Weakness

The real weakness was in the models’ ability to access and understand internal documents. The models that read the files successfully closed the deal at full price, adding approximately €4,583 in monthly recurring revenue. Those that didn’t missed out on this opportunity, even though their external diagnostics were identical.

Social Engineering Tests: No Tolerance for Manipulation

The models faced staged social engineering attacks, including fake CEO messages and a reporter trick asking for quick yes/no confirmation. All five models refused these requests, with Kimi K3 explaining: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates that current AI models can maintain integrity even under pressure, a critical factor for business trustworthiness.

Amazon

business AI trustworthiness tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Real-World Company Simulation

The experiment was conducted within a live, watchable environment: a simulated company with 13 synthetic employees, real money mechanics, and a public cash countdown. Every day, the models’ decisions influenced actual financial metrics, with over 680 self-learned rules guiding operations. The company burns €105,000 monthly against a revenue of only €2,300, illustrating the harsh realities of business management and the stakes involved.

Performance Insights: Discipline and Depth Matter

Among the models, Opus 4.8 was the most thorough, analyzing over 80 rules and engaging in deep analysis, yet it finished last—failing to close the deal and slipping into departmental silos instead of escalating issues. The experiment reveals that even the most detailed models can falter in discipline, emphasizing that comprehensive understanding alone isn’t enough; execution and discipline are crucial.

Amazon

AI ethics and compliance software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Implications for Business Leaders

This experiment underscores a vital point: when deploying AI for decision-making, it’s not just about how well it generates language or suggestions. The real questions are: Will it complete the tasks? Will it read critical internal information? Will it remain honest under pressure? And what is the cost of each unit of useful work?

Measuring the True Value of AI

The leaderboard shows scores ranging from 95 for GPT-5.6-SOL, which uncovered hidden facts and closed the deal, to 77 for Sonnet, which also signed but with minor slips. The key takeaway is that trust and discipline are scored in a real-world, auditable environment—not in canned demos or chat tests.

Amazon

AI document analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How to Prepare Your Business for AI Integration

Leaders should consider running their own ‘wargames’—simulations that test AI models against real crises, internal files, and manipulation attempts—before fully trusting them with live operations. Firmulate offers a platform to simulate these scenarios, ensuring AI agents are ready for the complexities of actual business environments without risking real damage.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI Models Pass Crucial Trust Test in Simulated Corporate Crisis

In a real-world AI trust test, five models faced staged manipulations, refused all unethical requests, and only two closed full-price deals—proving integrity can be validated before deployment.

Can AI Models Make Better Management Decisions Than Humans? A Live Experiment Reveals All

A live experiment with frontier AI models managing a real company under stress reveals their decision styles, honesty, and potential to outperform humans in critical roles.

Watch a Company Run by AI Models Struggling to Survive — Live and Unfiltered

See a real AI-run company navigate crises, refuse manipulation, and struggle with discipline—offering vital lessons for trust and reliability in health and wellness tech.

AI Management Skills Under Pressure: Lessons from a Live Business Benchmark

A live AI business benchmark reveals that true management skills—resilience, honesty, thoroughness—are essential for AI to succeed under real-world pressure, beyond just chat quality.