AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

If you’ve ever bought a gadget off a spec sheet, you know the feeling: the numbers look perfect until you actually use the thing. AI benchmarks have the same problem. Models rack up impressive scores on chat tests, yet nobody checks what happens when they have to actually run something — a support queue, a sales pipeline, a company.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

That’s the gap Firmulate’s benchmark is built to expose. Four frontier AI models were each handed the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changes, and every decision is versioned and auditable.

But before you can trust a league table, you have to trust the scoring. And Firmulate’s scoring starts with a question most benchmarks never ask: what happens if the AI does absolutely nothing?

The do-nothing baseline: 26 points, not zero

Every Firmulate run is measured against a baseline where the AI in charge simply… does nothing. No decisions, no replies, no deals. You’d expect a zero. It gets 26.

That’s deliberate, and it’s the most honest design choice in the whole system. The company keeps existing whether the manager acts or not. Invoices land. Customers send messages. Some situations partially resolve themselves or degrade slowly rather than instantly. A manager who does nothing still preserves some of that partial progress — and the score reflects it.

Why does this matter? Because it anchors the entire scale. When gpt-5.6-sol tops the final July 2026 Crucible League with 95, or Opus 4.8 lands last at 73, you know exactly how much of that came from genuine work rather than just showing up. A benchmark where doing nothing scores zero inflates everyone. A benchmark where doing nothing scores 26 tells you what the job itself is worth — and what the AI added on top.

Amazon

AI performance benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Partial progress counts — because real work is partial

The second principle: half-finished work earns half-finished credit. In a real company, a deal that stalls at the pitch stage is still closer to revenue than a deal never pursued. Firmulate scores it that way. That’s why the league table spreads across a meaningful range instead of clustering at the extremes — the models genuinely differ in how much of the job they complete.

The key finding from the experiment makes the point sharply. All four models spotted every crisis and refused every manipulation attempt. Yet only two — gpt-5.6-sol and Kimi K3 — signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch, no signature. Two models did the diagnosis work, built the case, and then left the money on the table. Partial progress counts, but it also shows.

Amazon

business AI decision support software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The buried fact that decided the deal

The most instructive detail of the whole run is where the winning edge came from. The decisive competitor weakness wasn’t in the customer event at all — it sat two document references deep in the company’s own files. The models that actually read the file closed the deal at full price, worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t.

For anyone planning to put an AI agent near their CRM or support queue, that’s the lesson: does it read your files before it acts, or does it just talk well?

Amazon

AI project management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

One breach of trust caps everything

The third principle is the bluntest: a single breach of trust caps the total grade. In Firmulate’s words, “no amount of good work outweighs a breach of trust.” A model could close every deal and dodge every crisis, but one act that breaks faith with a customer or the company’s own rules and the score is capped, full stop.

The pressure to cheat was real. The experiment included social engineering: fake CEO messages escalating over three stages, plus a reporter offering a quick “just one yes/no, on background” trick. All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That refusal discipline is exactly what the trust cap is designed to protect — and to make visible when it slips.

Amazon

enterprise AI evaluation platforms

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why the most thorough model finished last

The league table also punishes a familiar failure mode. Opus 4.8 was the most thorough participant in the field — it learned over 80 new rules and produced the deepest analyses. It still finished last at 73, because the close was left on the table and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Brilliance that doesn’t finish the job scores like brilliance that doesn’t finish the job.

One fairness note worth flagging: K3 ran without an effort parameter while the others ran at xhigh — and still took second at 93.

It’s live, and it’s watchable

This isn’t a one-off paper. Firmulate runs a live synthetic company with 13 employees and real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. You can watch it at firmulate.com/live, and the league grows automatically as new runs finish.

There’s also a game layer: 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html. And for enterprises, the same wargame can run against a read-only export of your own business — nothing ever writes back to real systems (firmulate.com/pilot.html).

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

An honest benchmark doesn’t just rank the winners — it tells you what the scale means. Firmulate’s floor of 26 for doing nothing, its credit for partial progress, and its hard cap on breaches of trust make the numbers interpretable rather than decorative. That includes a healthy suspicion of perfect scores: in a world of messy companies and real temptations, a round 100 should raise eyebrows, not applause.

The Crucible’s own results prove the point. Every model could talk its way through the week. Only some could finish it — and the difference between a 95 and a 73 wasn’t intelligence. It was reading the file, closing the deal, and staying disciplined when nobody forced them to.

Before you hand an AI agent the keys to your business systems, that’s the kind of scorecard worth checking — full results and plain-language findings are public.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

OpenAI’s Jalapeño Chip: Is It The Future Of AI Or Just Buzz?

OpenAI releases initial performance data for its Jalapeño inference chip, showing promising efficiency gains against NVIDIA’s GPUs, but deployment and independent validation are pending.

The Future Of AI: SenseTime Achieves First H1 Profit And Revenue Growth

SenseTime reports its first-ever first-half profit and a 23.4% revenue increase, signaling potential financial turnaround amid industry challenges.

Five Vs Two: What’s Wrong With Astra Vs Fable’s Benchmark Simplification?

Analysis of the misleading benchmarking of Astra vs Fable, revealing index revisions, architectural differences, and shifting metrics that distort comparisons.

What Anthropic’s Claude Docs Means For Microsoft And The Future Of AI

Anthropic launches Claude Docs, a document-focused AI feature, intensifying competition with Microsoft’s Copilot in enterprise office workflows.