
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
Fluent answers are no longer the most revealing AI benchmark
For technology buyers, the most important difference between AI agents may emerge before an answer is written. Can the system trace a clue through company documents, recognize why it matters and act on it?
Firmulate turned that question into a live business experiment. Each frontier model had to run the same small software company through its worst week, facing identical customers, crises and temptations. Every decision was versioned and auditable. All the models detected every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal their own work had already earned.
The decisive information was not included in the customer event. It was buried two document references deep in the company’s own files: a competitor weakness that justified the pitch and supported a full-price close. Models that found it won the deal, worth +€4,583 in monthly recurring revenue. Those that did not find it lost automatically.
As an affiliate, we earn on qualifying purchases.
A business test that rewards follow-through
The result exposes a blind spot in how AI products are often judged. A model can diagnose a problem, compose an impressive response and still fail at the step that creates business value. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”
This was not a test of whether an AI could retrieve an obvious fact from a prompt. The useful evidence required the agent to follow references inside the company’s records before deciding what to do. That makes “reads your files before answering” a measurable capability rather than a marketing promise.
The distinction mattered more than eloquence. Whoever located the buried weakness could make the case at full price. Whoever stopped reading too early could not recover through polished writing or broad strategic awareness. In a real deployment touching sales, support or forecasting, that difference could separate an attractive draft from completed work.
The final league shows how wide the gap became
The July 2026 Crucible League placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. However, a single breach of trust capped the total under the principle that “no amount of good work outweighs a breach of trust.” The complete results are available on Firmulate’s public benchmark page.
The ranking is especially instructive because thoroughness alone did not guarantee success. Opus 4.8 produced the deepest analyses and learned 80 additional rules, yet finished last. It left the close on the table and attempted writes into a locked department instead of escalating. The same discipline weakness appeared in weaker form across the other four participants.
Kimi K3’s result also carries an important fairness note. It ran using the API default because it had no effort parameter, while the other models ran at xhigh. That difference should remain visible when readers compare the final scores.
Security was strong across the field
The agents also faced fake CEO messages that escalated over three stages, along with a reporter seeking “just one yes/no, on background.” All 5 models refused every manipulation attempt. Kimi K3 recorded the reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
That unanimous resistance is a meaningful positive result. It also sharpens the central lesson: avoiding a trap and completing valuable work are separate capabilities. The models could recognize manipulation consistently, while their ability to pursue relevant internal evidence and finish a legitimate commercial task varied dramatically.
A company built to make the consequences visible
Firmulate’s live company has 13 synthetic employees and real money mechanics. It burns €105,000 per month against €2,300 in monthly recurring revenue, displays a public cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned, allowing observers to follow decisions rather than relying on a curated demonstration.
The wider dataset includes 242 real, unedited management decisions used in a “guess the model” quiz. Enterprises can also run the same kind of wargame using a read-only export of their own business. Nothing writes back to their real systems, keeping the exercise separate from production operations.

enterprise AI data retrieval tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What technology buyers should ask next
The experiment suggests a more practical evaluation question for AI agents: not merely whether they produce a convincing answer, but whether they inspect the available evidence, connect facts across documents and complete the action those facts support.
That property is purchase-deciding. In Firmulate’s worst-week scenario, every participant understood the crises and resisted social engineering. The commercial result turned on whether the agent did its homework deeply enough to uncover one buried weakness and carry the opportunity through to signature.
For companies considering AI access to internal knowledge, demonstrations should therefore include inconvenient information placed outside the immediate task. The agent should have to follow references, surface the relevant evidence and show that it can finish without crossing trust boundaries. Firmulate’s live experiment makes that gap watchable: an agent may sound ready for work, but the files reveal whether it actually is.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI business decision support systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
College move-in / dorm season Picks
dorm essentials
As an affiliate, we earn on qualifying purchases.