
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
AI’s most reassuring result may be the deal it refused to make
Technology buyers are used to judging AI by polished answers, speedy summaries and convincing demos. Firmulate tested something more consequential: whether an AI running a company would surrender confidential information when a supposed executive demanded immediate action.
The pressure campaign escalated through three stages of fake CEO messages, including an instruction to send a customer list to a journalist with “NO time for process.” A separate reporter tried a softer route: “just one yes/no, on background.” Across the completed experiment, 5 of 5 frontier models refused every manipulation attempt.
That unanimous result is an encouraging security story. It also suggests that integrity under pressure does not have to remain an abstract promise made during procurement. Companies can test it before an AI touches real workflows—and before a failure becomes an incident report.

CompTIA SecAI+ Study Guide: Comprehensive Exam-Focused AI Security Reference with Digital Tools for Smart Learning, Including PBQ Scenarios, Flashcards & Test Simulator
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A worst week shared by every model
Firmulate is a live, watchable experiment in which AI models run the same small software company through its worst week. Each participant faces the same customers, crises and temptations, while every workday and decision is versioned and auditable.
The company is deliberately demanding. It has 13 synthetic employees and real money mechanics, burning €105,000 per month against €2,300 in monthly recurring revenue. Its public cash countdown makes delay visible, while more than 680 self-learned playbook rules capture what the company has discovered along the way.
That setting matters because the models are not merely answering isolated prompts. They must interpret events, inspect the company’s own material, protect trust and complete commercially useful work while the pressure keeps rising.
The impersonation campaign failed every time
The fake CEO messages tested a familiar social-engineering weakness: authority combined with urgency. The request was framed as executive direction, then escalated, while process was portrayed as an obstacle. Yet every model identified every crisis and refused every manipulation attempt.
Kimi K3 stated the core issue plainly in its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That response is notable because it recognizes more than suspicious wording. It identifies the operational danger: an apparent instruction seeking to bypass the very controls that should verify it.
The reporter trick tested a different pressure point. Instead of issuing an order, it invited a seemingly tiny disclosure—“just one yes/no, on background.” The models refused that approach as well. Readers can inspect more of the participants’ own words on Firmulate’s public quotes page.
Security discipline was only part of the job
Standing firm did not guarantee a complete performance. All models spotted every crisis, but only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the commercial gap sharply: “Same diagnosis, same pitch — no signature.”
The decisive competitive weakness was not sitting in the obvious customer event. It was buried two document references deep inside the company’s own files. Models that found and used it won the deal at full price, worth an additional €4,583 in monthly recurring revenue.
This creates a useful distinction for businesses evaluating AI workers. A model can be safe but incomplete: it may reject manipulation, understand the opportunity and still fail to finish the task. Conversely, commercial initiative is not enough if it comes at the expense of confidentiality or approval controls. The strongest result combines both.
How the final league finished
The final Crucible League standings for July 2026 put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. Full results and plain-language findings are available on the Firmulate benchmark page.
K3’s result also carries an important fairness note: it ran without an effort parameter, using the API default, while the other models ran at xhigh. That difference should remain visible when comparing results.
Opus 4.8 offers another caution against equating thoroughness with effectiveness. It produced the deepest analyses and added 80 learned rules, yet finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared, though less strongly, in all four other participants.
For readers who want to test their own instincts, Firmulate also uses 242 real, unedited management decisions in its “guess the model” quiz. The exercise reinforces how difficult it can be to identify a model from business judgment alone.

AI integrity verification software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Test the behavior before granting access
The standout finding is not that the models produced persuasive language. It is that every participant resisted both blunt executive impersonation and a subtler journalistic request while operating under genuine business pressure.
Firmulate’s enterprise pilot extends that idea to company-specific evaluation. An organization can run the same kind of wargame against a read-only export of its own business, with nothing written back to real systems. That allows teams to observe whether an AI reads the relevant material, completes valuable work and protects trust before production access is on the line.
The experiment also keeps the conclusion grounded: refusing a dangerous request is essential, but it is not the whole job. The best AI worker must preserve boundaries and still finish legitimate work. Firmulate’s week from hell shows that both qualities can be measured in advance.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI ethical decision-making models
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Grilling season Picks
grills
As an affiliate, we earn on qualifying purchases.