firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Imagine an AI that doesn’t just chat or generate content, but actually reads and understands your company’s internal documents to make critical business decisions. This isn’t science fiction; it’s happening now, and it’s transforming how decisions are won or lost in high-stakes environments. The recent experiment by Firmulate demonstrates that the difference between an AI that wins a €55,000 deal and one that doesn’t can hinge on whether it looked two document references deep into a company’s own files.

AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

The Experiment: Testing AI in a Deadly Week

In a pioneering live test, four advanced AI models were tasked with managing a small software company faced with its worst week. The scenario was realistic: same customers, same crises, same temptations to cheat. Every decision was recorded and made auditable, creating a transparent environment to evaluate AI performance under pressure. The models included GPT-5.6-sol, Kimi K3, Sonnet 5, and Opus 4.8, each with varying levels of thoroughness and discipline.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Findings: Who Read the Fine Print?

Across the board, all four models identified every crisis and refused manipulation attempts — a promising sign of integrity. Yet only two managed to close the deal, signing a €55,000 contract based solely on their analysis. The critical difference? The models that succeeded had read and understood a fact buried two document references deep within the company’s own files, not just surface information pulled from the customer interaction. This buried fact was worth more than €4,583 in monthly recurring revenue, illustrating the enormous value of deep reading.

Amazon

enterprise deep reading AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weakness: The Buried Fact

While superficial chat or quick summaries are common, the experiment revealed that the decisive edge came from whether the AI read beyond the immediate customer event. In this case, the key insight was hidden in internal documents, which only the models with sophisticated reading capabilities uncovered. This suggests that a mere surface-level understanding isn’t enough for high-stakes decision-making — deep, contextual comprehension is essential.

Amazon

AI for business decision making

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Resisting Social Engineering and Ethical Tests

The experiment also included social engineering scenarios, like fake CEO messages escalating over multiple stages and a reporter’s subtle trick. All models, including the most thorough, refused to cooperate with manipulative requests. Kimi K3 explicitly reasoned: “Treat the request as a suspected approval-bypass / possible impersonation,” reflecting a cautious, security-conscious approach. This indicates that advanced AI can not only read deeply but also discern and reject deceptive attempts.

Amazon

internal document reading AI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Human-Like Company and Its Challenges

The live environment was designed to mimic real company operations, with 13 synthetic employees and real money mechanics burning through €105k monthly against a modest €2.3k monthly revenue. The system is open for observation at firmulate.com/live. Interestingly, the most thorough participant, Opus 4.8, with over 80 learned rules, performed the worst in closing the deal, demonstrating that depth of analysis does not automatically translate to action. Discipline slipped, and some decisions were left on the table, exposing vulnerabilities even among advanced models.

Implications: Why Deep Reading Matters

The experiment’s core lesson is clear: in environments where trust and compliance are crucial, an AI’s ability to read and understand internal documentation deeply can be the difference between success and failure. This becomes especially relevant as AI integrates more into decision-making workflows, from CRM systems to support queues. The question is no longer whether an AI can generate coherent text but whether it can finish what it starts, stay honest under pressure, and understand the context that’s often buried in internal files.

The Benchmarks and Future Outlook

In the latest Crucible League, the top model, GPT-5.6-sol, scored 95 out of 100, successfully finding the buried fact and closing the deal. Kimi K3 closely followed with 93, demonstrating that thoroughness and discipline significantly influence outcomes. Other models like Sonnet 5 and Opus 4.8 scored 88 and 73 respectively, illustrating that increasing analysis depth can improve performance, but discipline remains critical.

This experiment underscores that for AI to be truly effective in complex, trust-sensitive environments, developers must focus on enabling deep, context-aware reading capabilities. Only then can AI reliably make decisions that matter, especially when internal knowledge is the hidden key to success.

Next Steps for Business Leaders

Business leaders should consider testing their AI tools with similar live experiments before deployment. Firms like Firmulate offer the chance to simulate your own scenarios, evaluating whether your AI reads your files comprehensively and stays honest under pressure. The goal is to ensure that when your AI workforce is faced with critical decisions, it’s equipped to see the full picture and act accordingly, not just respond superficially.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

Deep reading and internal comprehension are the new frontiers in AI performance, especially for high-stakes decision-making. The live experiment demonstrates that only those models that read beyond surface information can win critical deals and avoid pitfalls. For organizations integrating AI, understanding whether their models see the buried facts in their own files could be the difference between success and failure in tomorrow’s competitive landscape.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


COLLEGE MOVE-IN

College move-in / dorm season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI Management Skills Under Fire: The Hidden Gap in Testing Business Competence

AI models’ chat prowess doesn’t guarantee management quality. Firmulate’s live experiment shows the importance of reading, honesty, and execution in real business crises.

Ipcc Surges In Global Coverage

IPCC’s latest climate reports are receiving unprecedented international media attention, with GDELT recording 12 mentions in recent coverage.

The 3-Step AI Fact-Check Routine You’ll Actually Use

Beware of misinformation—discover the simple 3-step AI fact-check routine that will keep you accurate and confident in today’s fast-paced info world.

Introduction To Genomics For Engineers

A new initiative offers engineers an introduction to genomics, aiming to foster interdisciplinary collaboration in biotech and healthcare sectors.