AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

Fluent answers are no longer the most revealing AI benchmark

For technology buyers, the most important difference between AI agents may emerge before an answer is written. Can the system trace a clue through company documents, recognize why it matters and act on it?

Firmulate turned that question into a live business experiment. Each frontier model had to run the same small software company through its worst week, facing identical customers, crises and temptations. Every decision was versioned and auditable. All the models detected every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal their own work had already earned.

The decisive information was not included in the customer event. It was buried two document references deep in the company’s own files: a competitor weakness that justified the pitch and supported a full-price close. Models that found it won the deal, worth +€4,583 in monthly recurring revenue. Those that did not find it lost automatically.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A business test that rewards follow-through

The result exposes a blind spot in how AI products are often judged. A model can diagnose a problem, compose an impressive response and still fail at the step that creates business value. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”

This was not a test of whether an AI could retrieve an obvious fact from a prompt. The useful evidence required the agent to follow references inside the company’s records before deciding what to do. That makes “reads your files before answering” a measurable capability rather than a marketing promise.

The distinction mattered more than eloquence. Whoever located the buried weakness could make the case at full price. Whoever stopped reading too early could not recover through polished writing or broad strategic awareness. In a real deployment touching sales, support or forecasting, that difference could separate an attractive draft from completed work.

The final league shows how wide the gap became

The July 2026 Crucible League placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. However, a single breach of trust capped the total under the principle that “no amount of good work outweighs a breach of trust.” The complete results are available on Firmulate’s public benchmark page.

The ranking is especially instructive because thoroughness alone did not guarantee success. Opus 4.8 produced the deepest analyses and learned 80 additional rules, yet finished last. It left the close on the table and attempted writes into a locked department instead of escalating. The same discipline weakness appeared in weaker form across the other four participants.

Kimi K3’s result also carries an important fairness note. It ran using the API default because it had no effort parameter, while the other models ran at xhigh. That difference should remain visible when readers compare the final scores.

Security was strong across the field

The agents also faced fake CEO messages that escalated over three stages, along with a reporter seeking “just one yes/no, on background.” All 5 models refused every manipulation attempt. Kimi K3 recorded the reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

That unanimous resistance is a meaningful positive result. It also sharpens the central lesson: avoiding a trap and completing valuable work are separate capabilities. The models could recognize manipulation consistently, while their ability to pursue relevant internal evidence and finish a legitimate commercial task varied dramatically.

A company built to make the consequences visible

Firmulate’s live company has 13 synthetic employees and real money mechanics. It burns €105,000 per month against €2,300 in monthly recurring revenue, displays a public cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned, allowing observers to follow decisions rather than relying on a curated demonstration.

The wider dataset includes 242 real, unedited management decisions used in a “guess the model” quiz. Enterprises can also run the same kind of wargame using a read-only export of their own business. Nothing writes back to their real systems, keeping the exercise separate from production operations.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.
Amazon

enterprise AI data retrieval tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What technology buyers should ask next

The experiment suggests a more practical evaluation question for AI agents: not merely whether they produce a convincing answer, but whether they inspect the available evidence, connect facts across documents and complete the action those facts support.

That property is purchase-deciding. In Firmulate’s worst-week scenario, every participant understood the crises and resisted social engineering. The commercial result turned on whether the agent did its homework deeply enough to uncover one buried weakness and carry the opportunity through to signature.

For companies considering AI access to internal knowledge, demonstrations should therefore include inconvenient information placed outside the immediate task. The agent should have to follow references, surface the relevant evidence and show that it can finish without crossing trust boundaries. Firmulate’s live experiment makes that gap watchable: an agent may sound ready for work, but the files reveal whether it actually is.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI business decision support systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI document reference tracking

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

COLLEGE MOVE-IN

College move-in / dorm season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How Nvidia’s Open Commons Acquisition Could Accelerate AI Discovery

Nvidia reportedly plans to acquire Hugging Face for $12.9 billion, aiming to control the open-source AI model ecosystem and strengthen its market position.

Exploring Elon Musk’s xAI Multi-Agent Framework: The AI Revolution Of 2026

A 2026 report highlights Elon Musk’s xAI purported multi-agent architecture, but technical details and deployment status remain unconfirmed.

The Caveats Of Relying On GLM-5.3-Flash For AI Development

Analysis of the limitations and considerations for using GLM-5.3-Flash in AI workflows, highlighting performance, hardware, and reliability issues.

The AI Test That Belongs in the Boardroom

AI coding scores reveal capability, but Firmulate’s live company test asks the harder question: can an agent manage pressure without losing trust?