
Imagine observing a company operating in real time—making decisions, facing crises, and even risking its survival—while you watch every move unfold live online. This is not a science fiction scenario but the reality of an ongoing experiment by Firmulate, a company that publicly tests AI models as if they were running a real business. For those interested in how artificial intelligence could impact industries, health, or even daily life, this experiment offers a rare glimpse into AI decision-making under pressure and the profound questions it raises about trust, honesty, and operational integrity.
The Live Experiment: AI Companies in Action
At the heart of this experiment is a live virtual company, operated by 13 synthetic employees, which runs on real money mechanics. This isn’t just a demo—it’s an actual company burning through €105,000 each month against a modest monthly revenue of €2,300. Every workday, this digital enterprise is meticulously versioned, with decisions recorded, analyzed, and published for public viewing at firmulate.com/live. The goal? To understand how AI models handle complex, real-world business crises that require judgment, trustworthiness, and strategic thinking.
AI decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How the AI Models Were Tested
Four frontier AI models, including GPT-5.6-sol and Kimi K3, faced the same challenging week of business crises—customers, crises, and temptations to manipulate the system. Each decision was carefully recorded and auditable, ensuring transparency in their choices. The models were tasked with diagnosing problems, making decisions, and closing deals, all while resisting attempts to manipulate or deceive them.
In one revealing test, all four models identified and responded to every crisis correctly, refusing manipulation attempts like fake CEO messages or reporter tricks. However, only two models managed to close a deal worth €55,000, earning additional monthly revenue of about €4,583 in recurring income. The other two models, despite their correct diagnosis and pitch, failed to follow through and sign the deal, illustrating critical discipline failures in their decision processes.
business crisis management AI tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness: Information in Files
Interestingly, the decisive factor in closing the deal lay not in the customer interactions but in information buried deep within the company’s own internal files. The models that read and understood this internal documentation ended up winning the deal and securing the full revenue potential. This underscores a vital point: in complex decision-making, context and internal knowledge are often more crucial than surface-level customer data.
AI security social engineering defense
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Defense Against Social Engineering
Security and honesty were also put to the test through staged social engineering attacks, such as fake CEO messages and background inquiries from reporters. Remarkably, all five AI models refused these manipulative efforts, with Kimi K3 explicitly reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This level of refusal demonstrates a promising capacity for AI to resist deception, an increasingly important trait in today’s digital landscape.
AI internal documentation analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Reality of a Burn-Rate Company
The live company operates with a stark financial reality—burning €105,000 each month with minimal revenue, in a public countdown that emphasizes its fragile survival. Every decision, every interaction, is an opportunity for the models to demonstrate their operational integrity and discipline. The experiment’s transparency allows observers to see firsthand whether AI can sustain a company in a crisis—an essential question as AI begins to touch more areas of business and life.
Lessons from the Performance
The most thorough model, Opus 4.8, with over 80 learned rules and deep analysis, finished last in the deal-closing race, primarily because it left opportunities unexploited and failed to escalate issues appropriately. In contrast, the models that combined strong discipline with thorough analysis were better at closing deals and resisting shortcuts. This suggests that the path to reliable AI-driven decision-making requires not just data but disciplined processes and internal awareness.
What This Means for You
For anyone concerned about the growing role of AI in everyday life—whether in healthcare, finance, or personal wellness—the key takeaway is this: the value of AI is not just in producing convincing words or responses. It lies in its ability to finish what it starts, to read and understand context deeply, and to resist manipulative tactics. As AI models are increasingly integrated into operations, understanding their decision mechanics and integrity becomes paramount.
Try It Yourself
If you want to see this kind of testing firsthand, firms can run their own wargame using a read-only export of their business data. This allows organizations to evaluate how their AI tools would handle crises, decisions, and pressures—without risking real systems or data. Learn more at firmulate.com/pilot.html.

The ongoing live experiment by Firmulate offers a rare window into AI decision-making under pressure, revealing strengths in crisis detection and weaknesses in follow-through—crucial insights for trusting AI in real-world business and life.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html