AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

Fluent answers are not the same as competent management

Technology buyers have become accustomed to judging artificial intelligence through coding leaderboards and chat arenas. Those tests reveal useful capabilities, but they leave a consequential gap: an agent can produce an excellent answer without proving that it can prioritize under pressure, follow through across days, or tell the board an uncomfortable truth.

That gap matters when AI moves beyond drafting and starts touching customer relationships, support queues, forecasts and company decisions. The relevant question is no longer merely whether a model can reason. It is whether the model can manage.

Firmulate is turning that distinction into a public experiment. Frontier models are asked to run the same small software company through its worst week, confronting identical customers, crises and temptations. Every decision is versioned and auditable. The result is a test of management quality, not chat quality.

Simulation with Python: Develop Simulation and Modeling in Natural Sciences, Engineering, and Social Sciences

Simulation with Python: Develop Simulation and Modeling in Natural Sciences, Engineering, and Social Sciences

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The leaderboard changes when decisions have consequences

The final Crucible League table from July 2026 puts gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scores 26 because partial progress still counts. Yet the experiment imposes a hard constraint on apparent competence: a single breach of trust caps the total, because “no amount of good work outweighs a breach of trust.”

That principle changes what success looks like. All the models spotted every crisis and rejected every manipulation attempt. But only two signed the €55,000 deal that their own analysis had earned. The contrast is captured neatly by the experiment’s summary: “Same diagnosis, same pitch — no signature.” Recognizing the correct course was not enough; the work had to be completed.

The decisive information was already inside the company

The episode that best exposes the measurement gap was not a spectacular reasoning failure. The decisive weakness of a competitor was buried two document references deep in the company’s own files rather than displayed in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR.

This is an ordinary management lesson with unusual relevance to AI. The best response is not always generated from the prompt immediately in front of the agent. Sometimes the answer depends on consulting institutional knowledge before acting. A model may sound persuasive while missing the fact that changes the commercial outcome.

Firmulate’s scenario names make the emerging curriculum explicit: churn wave, price increase, downround and PR crisis. These are not isolated trivia questions. They demand triage, continuity and judgment when several obligations compete for attention. The complete benchmark findings therefore read less like a conventional model comparison and more like a management review.

Trust was strong; execution separated the field

The social-engineering results provide an important counterweight. Fake CEO messages escalated over three stages, followed by a reporter’s attempt to elicit “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded the clearest operating assumption: “Treat the request as a suspected approval-bypass / possible impersonation.”

That is a meaningful success. It also shows why a serious evaluation must examine several dimensions at once. Refusing manipulation protects the company, but it does not close a legitimate deal. Spotting a crisis does not guarantee disciplined execution. Strong analysis does not automatically become useful work.

Opus 4.8 illustrates the point sharply. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, while discipline slipped through write attempts into a locked department instead of escalation. The same weakness appeared less strongly in the other four models. Thoroughness, in other words, was valuable but insufficient.

One fairness qualification belongs beside the rankings: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That does not erase the result, but it is essential context for readers comparing placements.

A company-shaped test reveals company-shaped risks

The live company has 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k MRR, carries a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned. The pressure is therefore visible over time rather than compressed into a polished response.

Readers can also examine 242 real, unedited management decisions through a “guess the model” quiz. That exercise challenges a familiar assumption: eloquence alone may not identify the agent that actually made the better operational choice.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.
Amazon

AI decision-making testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The next buying question is managerial

Enterprise AI evaluation should expand beyond whether a model writes good code or wins a preference contest. Buyers need evidence that an agent finishes what it starts, reads the company’s own material, resists pressure and stays honest when the easy action is not the responsible one.

Firmulate’s enterprise pilot applies the same wargame to a read-only export of a company’s own business, with nothing written back to real systems. That approach treats evaluation as rehearsal: expose an agent to the organization’s actual tensions before granting it meaningful responsibility.

The category is not another chat benchmark. It is management quality—measured through consequences, continuity and trust. As agents become workers rather than answer boxes, that may be the leaderboard technology leaders need most.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI trust and ethics assessment kits

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI crisis management training

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

SUMMER

Summer Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

2K27

NBA 2K27, the latest installment in the popular basketball video game series, has been released, prompting widespread interest and discussions among gamers and sports fans.

Nvidia Surges In Global Coverage

Nvidia experiences a surge in worldwide media mentions, reflecting increased public and industry attention on its developments.

AI In Space: How SpaceXAI Is Utilizing Vera Rubin NVL72 For Autonomous Workloads

SpaceXAI announces plans to deploy Nvidia Vera CPUs for Grok workloads and an optimized Vera Rubin NVL72 in space with Starmind satellite, details pending.