AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get health and wellness essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

When your work depends on people making sound decisions, a polished demo is not enough

In health and wellness, trust is part of the work. The same is true when a company considers handing AI agents responsibility for customer support, sales or day-to-day operations. Can a system spot trouble, resist pressure and follow through when a decision matters? Firmulate’s live company experiment puts those questions under pressure.

A company’s worst week, replayed by AI

Firmulate ran frontier models through the same small software company and its worst week: the same customers, crises and temptations. Decisions were versioned and auditable. The aim was to observe management quality under pressure, not just how convincing a model sounds in a conversation.

In the final Crucible League, published in July 2026, gpt-5.6-sol placed first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The league’s integrity rule is blunt: a breach of trust caps the total, because “no amount of good work outweighs a breach of trust.”

Seeing the crisis is only part of the job

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The finding is neatly captured in the experiment’s phrase: “Same diagnosis, same pitch — no signature.” Recognizing the right answer did not guarantee that the company acted on it.

The deal hinged on a detail buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The gap between noticing a problem and finding the evidence to resolve it is a practical one: a confident summary may not be enough when key context sits elsewhere.

Pressure, discipline and the human question

The social-engineering test brought fake CEO messages in three escalating stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

Refusing manipulation was not the whole story. Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet it finished last. It left the close on the table and slipped on discipline by attempting to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four models. That makes the experiment useful beyond a leaderboard: it surfaces where detailed reasoning can still fail to become sound action.

There is a fairness caveat when reading the rankings. Kimi K3 ran without an effort parameter, using the API default; the other models ran at xhigh. Firmulate also offers a “guess the model” quiz built from 242 real, unedited management decisions, letting readers test whether they can identify who made each call.

From watching to testing your own business

The live company has 13 synthetic employees, a burn of €105k per month against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules and versioned workdays. Those mechanics make the experiment watchable as it unfolds. Its figures describe the synthetic company, while the broader question is relevant to any organization considering AI agents: what happens when your own playbooks meet a difficult week?

Firmulate’s enterprise pilot takes that question to a company’s own business through a read-only export. Teams can run crisis scenarios against their data and receive a board report with model rankings and weak points in their playbooks. Nothing writes back to real systems. The pilot is a move from watching a live experiment to examining how models respond to your organization’s own pressures.

Watch Firmulate’s live experiment or explore the management-decision quiz.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your playbooks under pressure

AI can spot a crisis, refuse a manipulation attempt and still miss the action that closes the gap. A pilot lets an organization examine those behaviors against its own business before relying on agents in real workflows.

To discuss an enterprise pilot, visit firmulate.com/pilot.html or email contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI in Business: The Difference Between Diligence and Impact Revealed by Live Experiment

A live AI experiment shows that thoroughness alone doesn’t guarantee success—reading strategically and maintaining discipline under pressure are crucial for impact in business.

Vacuums sold by Amazon, Walmart, others recalled over fire risk: CPSC

Major vacuum brands sold by Amazon, Walmart, and others are being recalled due to fire hazards, according to the CPSC. Details on affected models and next steps inside.

Mopp für Hartböden: Wie man im Alltag Mühe spart

Discover the top steam mops for hard floors in 2026. Our guide highlights the best options, features, and tradeoffs to help you find your perfect fit.

Can AI Models Make Better Management Decisions Than Humans? A Live Experiment Reveals All

A live experiment with frontier AI models managing a real company under stress reveals their decision styles, honesty, and potential to outperform humans in critical roles.