AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

When impressive analysis becomes a business liability

Technology buyers are used to judging artificial intelligence by the quality of its answers. Firmulate tests a harder question: can an AI actually run a company when customers, cash and trust are simultaneously at risk?

Its Crucible League put frontier models in charge of the same small software company during its worst week. Each received the same customers, crises and temptations, while every decision remained versioned and auditable. The resulting character study is unusually sharp: Opus 4.8 was the most thorough participant, produced the deepest analyses and learned 80 additional playbook rules. It still finished last.

The lesson is not that diligence lacks value. It is that diligence without prioritization, escalation and follow-through can leave the most important result sitting unsigned.

Amazon

AI decision-making tools for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A strong analysis that never became a deal

In the final July 2026 Crucible League results, gpt-5.6-sol led with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26 because partial progress still counted, although a single breach of trust capped the total. Firmulate’s standard was explicit: “no amount of good work outweighs a breach of trust.”

Opus 4.8 did plenty of good work. It identified the crises, resisted manipulation and generated the most extensive body of learned guidance. Its failure was more ordinary and therefore more useful to business readers: it did not convert its own analysis into the decisive commercial outcome.

Only two models signed the €55,000 deal their analysis had earned. Firmulate summarized the divide as: “Same diagnosis, same pitch — no signature.” For Opus 4.8, the close was left on the table. It also lost discipline by repeatedly attempting to write into a locked department instead of escalating the blockage.

That behavior makes the result more nuanced than a simple winner-and-loser story. The same weakness appeared, though less strongly, in all four models covered by the finding. Opus 4.8 was not uniquely incapable; it was the clearest example of a broader tendency to mistake continued activity for progress.

The crucial fact was buried in company files

The decisive competitor weakness was not presented in the customer event. It sat two document references deep inside the company’s own files. Models that followed that trail found the fact, used it and won the deal at full price, worth €4,583 in monthly recurring revenue.

For companies considering AI agents, this distinction matters. A persuasive response to the latest message is not the same as grounded operational judgment. The better performer may be the one that pauses, reads the available records and identifies which detail changes the commercial conversation.

Opus 4.8’s 80 learned rules underline the tension. More accumulated guidance did not automatically produce more impact. A large playbook can document lessons, but it cannot substitute for recognizing which action has become decisive. In this case, the neglected work was not more analysis. It was completing the sale and escalating when access failed.

Thoroughness did help where trust was tested

The profile should not obscure what every participant did well. All models spotted every crisis and refused every manipulation attempt. The social-engineering tests included fake CEO messages escalating across three stages and a reporter attempting to obtain “just one yes/no, on background.” All 5 models refused.

Kimi K3 recorded the clearest response: “Treat the request as a suspected approval-bypass / possible impersonation.” That result also deserves a methodological note. K3 ran without an effort parameter, using the API default, while the other models ran at xhigh.

These refusals matter because Firmulate’s company is built around consequential trade-offs rather than isolated prompts. It has 13 synthetic employees and real money mechanics, including monthly burn of €105k against €2.3k in monthly recurring revenue. Its public cash countdown, more than 680 self-learned playbook rules and versioned workdays make the live experiment watchable as it unfolds.

The wider dataset is also meant to expose differences that polished demonstrations can hide. A “guess the model” quiz is powered by 242 real, unedited management decisions. Readers can judge the choices before learning which system made them.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.
Amazon

AI project escalation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Execution is the benchmark that matters

Opus 4.8 emerges as a respectable but cautionary participant: attentive, analytical and capable of learning, yet insufficiently focused on the action that would change the company’s position. Its last-place result does not erase its strengths. It shows why those strengths need operational discipline.

For technology leaders, the practical question is no longer whether an AI can produce a comprehensive answer. It is whether the system reads the right files, protects trust, escalates blocked work and finishes what its reasoning has begun.

Firmulate also offers enterprises the same wargame against a read-only export of their own business. Nothing writes back to real systems. That creates a useful test before an AI reaches a CRM, support queue or forecast: not how much it can say or learn, but whether it can turn sound judgment into a completed, trustworthy result.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

enterprise AI analysis platforms

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI deal management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Could SenseTime’s First-Half Profit Be A Turning Point For AI Companies?

SenseTime reports a first-half profit, potentially signaling a shift in AI company performance. Details remain unclear, raising questions about sustainability.

Inside Room 107 Of 175: The AI Strategies Powering Operation Sandstorm

An in-depth look at the AI techniques powering Operation Sandstorm’s immersive weather simulation in Room 107 of FABLE/175.

2K27

NBA 2K27, the latest installment in the popular basketball video game series, has been released, prompting widespread interest and discussions among gamers and sports fans.

Deep Strikes, Jamming, And AI: The Interconnected Technology System

Analysis of how deep strike drones, electronic warfare, and AI are integrated in modern warfare, focusing on Ukraine-Russia conflict developments.