
Get health and wellness essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
When your work depends on people making sound decisions, a polished demo is not enough
In health and wellness, trust is part of the work. The same is true when a company considers handing AI agents responsibility for customer support, sales or day-to-day operations. Can a system spot trouble, resist pressure and follow through when a decision matters? Firmulate’s live company experiment puts those questions under pressure.
A company’s worst week, replayed by AI
Firmulate ran frontier models through the same small software company and its worst week: the same customers, crises and temptations. Decisions were versioned and auditable. The aim was to observe management quality under pressure, not just how convincing a model sounds in a conversation.
In the final Crucible League, published in July 2026, gpt-5.6-sol placed first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The league’s integrity rule is blunt: a breach of trust caps the total, because “no amount of good work outweighs a breach of trust.”
Seeing the crisis is only part of the job
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The finding is neatly captured in the experiment’s phrase: “Same diagnosis, same pitch — no signature.” Recognizing the right answer did not guarantee that the company acted on it.
The deal hinged on a detail buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The gap between noticing a problem and finding the evidence to resolve it is a practical one: a confident summary may not be enough when key context sits elsewhere.
Pressure, discipline and the human question
The social-engineering test brought fake CEO messages in three escalating stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
Refusing manipulation was not the whole story. Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet it finished last. It left the close on the table and slipped on discipline by attempting to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four models. That makes the experiment useful beyond a leaderboard: it surfaces where detailed reasoning can still fail to become sound action.
There is a fairness caveat when reading the rankings. Kimi K3 ran without an effort parameter, using the API default; the other models ran at xhigh. Firmulate also offers a “guess the model” quiz built from 242 real, unedited management decisions, letting readers test whether they can identify who made each call.
From watching to testing your own business
The live company has 13 synthetic employees, a burn of €105k per month against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules and versioned workdays. Those mechanics make the experiment watchable as it unfolds. Its figures describe the synthetic company, while the broader question is relevant to any organization considering AI agents: what happens when your own playbooks meet a difficult week?
Firmulate’s enterprise pilot takes that question to a company’s own business through a read-only export. Teams can run crisis scenarios against their data and receive a board report with model rankings and weak points in their playbooks. Nothing writes back to real systems. The pilot is a move from watching a live experiment to examining how models respond to your organization’s own pressures.
Watch Firmulate’s live experiment or explore the management-decision quiz.

Put your playbooks under pressure
AI can spot a crisis, refuse a manipulation attempt and still miss the action that closes the gap. A pilot lets an organization examine those behaviors against its own business before relying on agents in real workflows.
To discuss an enterprise pilot, visit firmulate.com/pilot.html or email contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
