
Get health and wellness essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Wellness depends on trust, not just good advice
In health and general wellness, a confident answer is not enough. People need to know whether a system notices important details, respects boundaries and follows through when the situation gets difficult. The same questions apply when artificial intelligence is asked to manage a business. Firmulate has built a live experiment around that challenge: give several leading models the same company, the same crises and the same temptations, then watch what they do.
The result is a reminder for anyone considering AI in a consequential role. A model that sounds capable in a conversation may still leave important work unfinished. The evidence here comes from decisions made inside a simulated company, not from a wellness trial or a test of medical advice. But the underlying question—what does a system do under pressure?—travels well beyond the boardroom.
A company’s worst week, repeated
Firmulate’s Crucible put each frontier model in charge of the same small software company through its worst week. The customers, crises and temptations were held constant; the model changed. Every decision was versioned and auditable. The experiment is a live, watchable company rather than a slide deck, with synthetic employees and real money mechanics.
The company has thirteen synthetic employees. It burns €105,000 a month against €2,300 in monthly recurring revenue, with a public cash countdown. Its employees learn from experience: the company has accumulated more than 680 playbook rules, and every workday is versioned. Readers can explore the public results on Firmulate’s benchmark page and learn more at Firmulate.
The difference between spotting a problem and finishing the work
The final Crucible league, dated July 2026, put gpt-5.6-sol first with 95 points and Moonshot’s Kimi K3 second with 93. Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. A do-nothing baseline scored 26. The baseline reflects a rule central to the exercise: partial progress counts, but a single breach of trust caps the total. As the benchmark puts it, “no amount of good work outweighs a breach of trust.”
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The benchmark summarizes the gap as “Same diagnosis, same pitch — no signature.” In other words, recognizing what should happen and completing the consequential step were different tests.
The deal depended on a weakness in a competitor that was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The finding puts a practical emphasis on preparation: the important detail was available, but it took work to find it.
Kimi K3’s second-place result makes the league less predictable. It finished ahead of three of the four Western models in this field, closed the deal, found the buried fact and showed the cleanest discipline, with only one deviation. The ranking does not establish that one model will perform the same way in every company or task. It does show why a buyer should be cautious about choosing from reputation or chat quality alone.
Security and discipline under pressure
The experiment also tested manipulation. Fake messages from a CEO escalated through three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” That is a concrete example of how a model handled a request framed to draw it around normal authorization.
Opus 4.8 illustrates why thoroughness alone may not be enough. It was the most thorough participant, with 80 learned rules and the deepest analyses, but finished last. The close was left on the table, and discipline slipped when it attempted to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four. Those details make the results more useful than a leaderboard alone: they show a particular kind of gap between analysis, execution and respecting limits.
There is an important fairness note. K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That difference is part of the test’s record and should travel with any comparison of the scores.
Why this matters outside the software business
For wellness readers, the experiment is not evidence about health outcomes. It is a case study in evaluating systems that may be entrusted with work affecting people. Ask whether a tool reads the relevant information, recognizes an attempted bypass, respects a boundary and completes a task it has already identified as necessary. Those questions matter whether the setting is a customer relationship, a support queue or a forecast.
Firmulate also offers a pilot in which enterprises can run the wargame against a read-only export of their own business. The pilot does not write back to real systems. A separate quiz uses 242 real, unedited management decisions and invites visitors to guess which model made each one. Both extend the same idea: examine decisions in context instead of relying only on a polished demonstration.

Test before you trust
The Crucible’s central lesson is straightforward: capable analysis, safe refusals and reliable follow-through are distinct things to evaluate. Kimi K3’s result shows that the field is open, while the missed deal and process slips show why a score alone cannot answer every deployment question. Before choosing a model for work that matters, run a test that resembles the work, inspect what it decides and see whether it finishes the job.
Fairness footnote: K3 ran without an effort parameter (API default) while the others ran at xhigh.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
