AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
Live on firmulate.com.

Imagine trusting an AI to run your business — not just to answer questions, but to make decisions that keep the lights on and the deals flowing. The latest experiment from Firmulate puts this to the test, pitting four advanced AI models against the same week of crises in a real company. The results reveal a stark truth: chat demo scores are not enough. What truly measures an AI’s business readiness is its ability to execute, stay honest under pressure, and close deals at full value.

How do we test AI’s business skills?

In a groundbreaking live experiment, four frontier AI models — including the highly-rated gpt-5.6-sol and newcomer Kimi K3 — each managed the same small software company’s toughest week. This company faced genuine crises: customer issues, trust manipulations, and even social engineering attacks designed to test the AI’s integrity. Every decision by the models was versioned and auditable, mimicking real business operations rather than canned chat responses.

AI Builders: Making The Decisions That Turn AI Code Into Real Software

AI Builders: Making The Decisions That Turn AI Code Into Real Software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The surprising performance scores

  • The models were scored on their crisis detection, honesty, and decision quality, with the top score being 95 for gpt-5.6-sol.
  • Kimi K3 scored 93, showing a clean decision record and closing the deal.
  • Sonnet 5 followed closely at 88, also closing the deal but with minor slips.
  • Fable 5, although disciplined in following rules (77), left a crucial opportunity unexploited.

Interestingly, the baseline score was just 26, highlighting that partial progress or minor work doesn’t equate to real business value. The key difference was that only two models actually signed the €55,000 deal—their own analysis had earned it. The others, despite diagnosing the same problems, failed to execute the close.

Amazon

business AI automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Beyond chat: reading the hidden files

The secret to the successful models was a buried fact deep in the company’s files—information that wasn’t obvious in the initial crisis reports. Reading and understanding this internal data was the decisive advantage that led to winning the deal at full price (+€4,583 MRR). This illustrates a vital point: surface-level chat demos often overlook an AI’s ability to dig into core data, which is critical for real business decisions.

Amazon

AI data analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Resisting manipulation and social engineering

All four models refused social engineering tricks, such as staged CEO messages and a reporter’s fake approval requests. For instance, Kimi K3 explicitly treated the request as a suspected impersonation attempt, demonstrating how well-designed prompts can test an AI’s integrity in high-pressure scenarios.

Amazon

AI cybersecurity social engineering

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The real-world company testbed

The live company managed by the AI models comprised 13 synthetic employees with real money mechanics, burning €105k monthly against €2.3k in MRR. It was a public, watchable experiment, with every decision made during real operational hours and every rule learned and applied through over 680 self-generated playbook instructions. All this is accessible at firmulate.com/live.

The discipline gap: why the best chat models can still fail to close deals

OPUS 4.8, the most thorough participant with over 80 learned rules, showed that even deep analysis doesn’t guarantee success. It left the deal on the table and slipped into process slips, writing attempts into a locked department instead of escalating. This highlights a critical insight: knowing what to do isn’t enough—doing it under pressure is the real challenge.

The lesson for the AI business landscape

This experiment underscores an essential reality: evaluating AI solely by chat quality is misleading. The true test is whether an AI can finish tasks, read internal data, resist manipulation, and close deals at full value. For business leaders, the question isn’t just “can it write?” but “will it finish what it starts?” and “can it stay honest when stakes are high?”

Test your AI against your own business

If you’re considering deploying AI in your operations, firms can run similar live tests through Firmulate’s platform. This allows companies to simulate their specific crises and workflows in a safe, read-only environment, revealing whether an AI model can truly perform — not just chat.

Why this matters for everyone using AI

As AI models become integral to customer support, sales, and decision-making, their ability to execute reliably and ethically is paramount. The outcome of this real-world test shows that the true measure of AI readiness isn’t in how well it chats — it’s in whether it can deliver results, stay honest, and close at full value under pressure.

Infographic — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Leverage AI To Improve Student Organization Efficiency

New AI-powered tools are transforming student organization by improving scheduling, note-taking, and resource management. Here’s what’s confirmed so far.

Removing React.js From The Codebase And Adapting Htmx For UI Interactivity (2023)

A major tech project has removed React.js from its codebase, adopting Htmx for UI interactivity. The change aims to simplify architecture and improve performance.

Leading AI-Enabled Wireless Earbuds In 2026: The Top 10

Discover the leading AI-powered wireless earbuds of 2026, featuring the best in sound, noise cancellation, and smart features for every budget.

AI Adoption: A Slow But Steady Path To Lasting Change

Analysis of how enterprise AI adoption remains slow yet resilient, with incumbents maintaining dominance through deep integration and data moats.