
When we evaluate AI systems, especially those poised to run real businesses, our focus is often on how well they produce answers—be it code, chat, or recommendations. But do these models truly understand management under pressure? The latest experiment from Firmulate reveals a stark truth: excelling in chat does not mean excelling at leading a company through crises.
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
Behind the Curtain of AI Leadership Testing
Most AI benchmarks highlight answer accuracy or problem-solving skills, but they overlook a crucial aspect—how management quality manifests when a company faces its worst week. The Firmulate live experiment placed four advanced AI models into a simulated small software company undergoing a brutal week of crises, customer demands, and ethical temptations. The goal? To see if these models could manage real-world pressures, make honest decisions, and ultimately close a lucrative deal.
The Setup: A Week of Crises and Temptations
Every model was tasked with navigating identical scenarios—customer churn, price hikes, PR crises, and internal decisions—repeatedly faced by real companies. Every decision was meticulously versioned and auditable, and the models had access to the company’s own files, including critical buried information. The environment was designed to test beyond simple problem-solving: it examined resilience, integrity, and decision quality under duress.
What the Models Achieved
- All models identified every crisis and refused manipulative or deceptive tactics, showing they could recognize unethical pressures.
- Despite this, only two out of four actually secured and signed the €55,000 deal their analysis had earned—demonstrating that recognizing problems is not enough.
- The decisive factor? Deep document comprehension. The winning models read two layers into the company’s own files, uncovering hidden crucial facts that led to the deal’s success.
The Real Weakness: Reading and Trust
Interestingly, the models with the best performance didn’t just rely on superficial answers; they read and understood internal documents, which proved pivotal. Those that missed this buried information failed to close the deal, despite recognizing crises and refusing manipulations.
Behavior Under Social Engineering
The models faced staged social engineering attacks: staged CEO messages escalating in severity and even a reporter trick asking for a quick background approval. All models refused to comply, aligning with Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This shows a level of ethical guardrail management that aligns with human management standards.
AI management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Human-Like Failures of ‘Deep’ Models
Among the models, Opus 4.8 demonstrated the most thorough internal analysis, with over 80 learned rules and deep assessments. Yet, it still finished last—failing to escalate internally and leaving critical decisions on the table, which cost the company the deal. The pattern was consistent across models: thoroughness didn’t guarantee execution under pressure.
The League Table and What It Tells Us
In the final standings—based on the ability to identify, manage, and close the deal—the scores ranged from 95 to 73. The top scorer, gpt-5.6-sol, uncovered the buried fact and signed the deal, exemplifying comprehensive management and information processing. The others, despite high scores, failed to close or slipped in discipline, revealing the invisible gap in current testing approaches.
business crisis management training tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Business and AI Adoption
The experiment underscores a critical insight: standard chat-based benchmarks are insufficient. The real test of an AI’s readiness to manage a company isn’t just its ability to produce correct answers—it’s whether it can read relevant internal information, stay honest under pressure, and follow through on commitments.
As organizations increasingly integrate AI into their CRM, support, and forecasting tools, understanding these management qualities becomes essential. The question isn’t just “Can it write well?” but rather: “Will it finish what it starts? Will it read and understand your files? Will it stay honest under pressure?”
Practical Tools for Managers
Firmulate offers enterprises a way to run their own management wargames—against a read-only export of their business, not affecting actual systems. This allows companies to see how different models perform in their unique context, exposing hidden gaps before deploying AI in critical decision-making roles.
Visit Firmulate to see the live experiment in action, explore the full results, or try the quiz that reveals how well your AI management tools measure up in real-world scenarios.

AI’s ability to generate answers isn’t enough for management tasks. Real leadership under pressure depends on reading internal information, staying honest, and executing with discipline. Firmulate’s live experiment exposes these hidden gaps—vital insights for any organization integrating AI into decision-making.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
internal document analysis AI tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
ethical decision-making AI models
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.