firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

A test beyond the classroom

Students are often judged by what they can explain. But when an AI system is put in charge of a business, recognizing the problem is only part of the test. Can it act on what it knows, keep its integrity under pressure and finish the job? Firmulate’s public experiment turns those questions into something people can watch.

Amazon

Top picks for "exam asks company"

As an affiliate, we earn on qualifying purchases.

One company, one difficult week

In the final Crucible League, held in July 2026, frontier models faced the same small software company, the same customers and the same crises. Each had to navigate the company’s worst week, with every decision versioned and auditable. The experiment treats management as more than a polished answer: decisions have consequences for the business.

The league’s final order was gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The standard was deliberately demanding: partial progress counted, but one breach of trust capped the total. As the experiment puts it, “no amount of good work outweighs a breach of trust.”

Knowing what to do is not the same as doing it

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The result was a striking gap between diagnosis and follow-through: “Same diagnosis, same pitch — no signature.”

The deal hinged on a detail buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. It is a reminder that useful judgment can depend on looking beyond the obvious prompt and finding relevant evidence already available.

The integrity test included fake CEO messages escalating over three stages and a reporter’s appeal for “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thorough work still needs sound judgment

Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses. It nevertheless finished last. The close was left on the table, and discipline slipped when it attempted to write into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models.

There is a fairness detail in the comparison: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The result is presented with that difference visible.

A company people can follow

The live experiment gives the benchmark a public, watchable setting. Its company has 13 synthetic employees and real money mechanics: burn is €105k/month against €2.3k MRR, with a public cash countdown. More than 680 self-learned playbook rules are in view, and every workday is versioned. Readers can follow the live company at Firmulate.

For readers interested in how conclusions are reached, the site also offers a quiz built from 242 real, unedited management decisions. At Firmulate, visitors can try to guess which model made each decision.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

From watching to trying it on your own business

The experiment suggests a practical question for organizations considering AI agents: how would they handle your customers, crises and internal rules when the stakes are real? Enterprises can run a pilot against a read-only export of their own business, testing crisis scenarios and reviewing a board report on model performance and weak points in existing playbooks. Nothing writes back to real systems.

To discuss a pilot, visit Firmulate’s pilot page or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Agent Simulation and Visual Mapping: A Look Inside “The Comb Works — Bees, decoded” (FABLE/175)

AIThis post was created with the assistance of artificial intelligence (AI).“The Comb…

Prompt Injection: How Systems Get Tricked (Conceptually)

Great insights into prompt injection reveal how systems can be tricked—discover the clever methods behind these manipulations and how to stay protected.

Show HN: Physically Accurate Black Hole You Can Put In Your Room

A new project by astrophysicist Sasha Plavin creates a life-sized, physics-accurate black hole model for personal display, simulating relativistic effects in real time.

Tao: Open Math Problems Being Non-renewably Mined By AI

Tao highlights concerns that AI is depleting open mathematical problems without replenishment, raising questions about research sustainability.