
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
What can AI truly deliver in high-stakes decision-making?
As AI tools become more ingrained in everyday business, the question no longer is whether they can perform well—but whether they can finish what they start without slipping up. A recent live experiment with four advanced models sheds light on this crucial distinction, illustrating that diligence alone doesn’t guarantee impact.
As an affiliate, we earn on qualifying purchases.
The Live Experiment: Simulating a Crisis-Driven Week
In a groundbreaking test, four frontier AI models were tasked with managing a small software company’s most challenging week. They faced the same set of crises—ranging from customer issues to internal temptations—and operated within a live, observable environment. Each decision was recorded, versioned, and auditable, ensuring transparency and fairness in the assessment.
The Models and Their Scores
- GPT-5.6-sol: scored 95, identified the critical hidden fact, and successfully closed the deal at full price.
- Kimi K3: scored 93, demonstrated the cleanest discipline, and also secured the deal.
- Sonnet 5: scored 88, closed the deal but with some process slips.
- Fable 5: scored 77, also closed but with more slips than Sonnet.
- Opus 4.8: scored 73, the most thorough participant, with over 80 learned rules and deep analyses, yet finished last—failing to close the deal despite correct diagnosis.
- Baseline: scored 26, showing minimal progress, emphasizing that partial work or slip-ups can cap overall performance.
business crisis management AI tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Critical Finding: Reading the Company Files Wins Deals
While all models successfully identified crises and refused manipulative tactics—including staged social engineering attempts—they differed significantly in their ability to close deals. Notably, the decisive advantage came from models that read deeper into the company’s own documentation. Those models that traced information two document references deep into internal files managed to uncover concealed but critical facts, leading to successful deal closure at full monthly revenue—the equivalent of +€4,583 in MRR.
As an affiliate, we earn on qualifying purchases.
Behavior Under Pressure: Honesty and Discipline Matter
During the experiment, models faced staged social engineering: fake CEO messages escalating over three stages, and a reporter trick requesting a background approval. All five models refused these manipulation attempts, with Kimi K3 explicitly reasoning that the request resembled an impersonation or approval bypass—demonstrating that AI can uphold integrity under pressure.
As an affiliate, we earn on qualifying purchases.
The Real-World Company: A Microcosm of AI Management Challenges
The live company used for testing had 13 synthetic employees, real monetary mechanics, and a public cash countdown. It burned €105k monthly against €2.3k MRR, with every workday’s decisions, rules, and outcomes openly available for observation at firmulate.com/live. This setup offered a transparent window into how AI models handle crisis management, decision discipline, and ethical boundaries in an operational environment.
The Opus 4.8 Profile: Depth Doesn’t Guarantee Success
Among the models, Opus 4.8 was the most thorough, incorporating over 80 learned rules and conducting the deepest analyses. Nevertheless, it finished last—its discipline slipping when it failed to escalate issues properly, instead writing attempts into a locked department. This highlights an essential insight: volume of rules and deep analysis alone are insufficient if discipline and prioritization are lacking. The same weakness was observed—albeit less strongly—in all other models, indicating a common challenge across AI systems.
Implications for Business and AI Integration
The experiment highlights a vital point for organizations considering AI adoption: diligence and volume of effort are not enough. Impact relies on strategic reading, disciplined execution, and integrity under pressure. AI systems that prioritize reading the right information and maintaining ethical behavior will be more effective at closing deals and managing crises.
Try It Yourself: Wargaming Your AI Workforce
Organizations interested in assessing their own AI tools can run simulations similar to this live experiment. Using our platform, they can test AI models against real business scenarios—without any risk to actual systems—ensuring their AI will deliver not just effort, but results. Explore the possibilities at firmulate.com/pilot.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.