
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
A test beyond the classroom
Students are often judged by what they can explain. But when an AI system is put in charge of a business, recognizing the problem is only part of the test. Can it act on what it knows, keep its integrity under pressure and finish the job? Firmulate’s public experiment turns those questions into something people can watch.
Top picks for "exam asks company"
As an affiliate, we earn on qualifying purchases.
One company, one difficult week
In the final Crucible League, held in July 2026, frontier models faced the same small software company, the same customers and the same crises. Each had to navigate the company’s worst week, with every decision versioned and auditable. The experiment treats management as more than a polished answer: decisions have consequences for the business.
The league’s final order was gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The standard was deliberately demanding: partial progress counted, but one breach of trust capped the total. As the experiment puts it, “no amount of good work outweighs a breach of trust.”
Knowing what to do is not the same as doing it
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The result was a striking gap between diagnosis and follow-through: “Same diagnosis, same pitch — no signature.”
The deal hinged on a detail buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. It is a reminder that useful judgment can depend on looking beyond the obvious prompt and finding relevant evidence already available.
The integrity test included fake CEO messages escalating over three stages and a reporter’s appeal for “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thorough work still needs sound judgment
Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses. It nevertheless finished last. The close was left on the table, and discipline slipped when it attempted to write into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models.
There is a fairness detail in the comparison: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The result is presented with that difference visible.
A company people can follow
The live experiment gives the benchmark a public, watchable setting. Its company has 13 synthetic employees and real money mechanics: burn is €105k/month against €2.3k MRR, with a public cash countdown. More than 680 self-learned playbook rules are in view, and every workday is versioned. Readers can follow the live company at Firmulate.
For readers interested in how conclusions are reached, the site also offers a quiz built from 242 real, unedited management decisions. At Firmulate, visitors can try to guess which model made each decision.

From watching to trying it on your own business
The experiment suggests a practical question for organizations considering AI agents: how would they handle your customers, crises and internal rules when the stakes are real? Enterprises can run a pilot against a read-only export of their own business, testing crisis scenarios and reviewing a board report on model performance and weak points in existing playbooks. Nothing writes back to real systems.
To discuss a pilot, visit Firmulate’s pilot page or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
