
Imagine trusting an AI to run your business — not just to answer questions, but to make decisions that keep the lights on and the deals flowing. The latest experiment from Firmulate puts this to the test, pitting four advanced AI models against the same week of crises in a real company. The results reveal a stark truth: chat demo scores are not enough. What truly measures an AI’s business readiness is its ability to execute, stay honest under pressure, and close deals at full value.
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
How do we test AI’s business skills?
In a groundbreaking live experiment, four frontier AI models — including the highly-rated gpt-5.6-sol and newcomer Kimi K3 — each managed the same small software company’s toughest week. This company faced genuine crises: customer issues, trust manipulations, and even social engineering attacks designed to test the AI’s integrity. Every decision by the models was versioned and auditable, mimicking real business operations rather than canned chat responses.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The surprising performance scores
- The models were scored on their crisis detection, honesty, and decision quality, with the top score being 95 for gpt-5.6-sol.
- Kimi K3 scored 93, showing a clean decision record and closing the deal.
- Sonnet 5 followed closely at 88, also closing the deal but with minor slips.
- Fable 5, although disciplined in following rules (77), left a crucial opportunity unexploited.
Interestingly, the baseline score was just 26, highlighting that partial progress or minor work doesn’t equate to real business value. The key difference was that only two models actually signed the €55,000 deal—their own analysis had earned it. The others, despite diagnosing the same problems, failed to execute the close.

Master AI for Beginners: Develop Artificial Intelligence Basics, Understand Machine Learning, and Unlock the Power of Automation for Business Productivity, and Everyday Life (The AI Success Series)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Beyond chat: reading the hidden files
The secret to the successful models was a buried fact deep in the company’s files—information that wasn’t obvious in the initial crisis reports. Reading and understanding this internal data was the decisive advantage that led to winning the deal at full price (+€4,583 MRR). This illustrates a vital point: surface-level chat demos often overlook an AI’s ability to dig into core data, which is critical for real business decisions.

Practical AI Security: A Hands-on Guide to Attacking, Defending, and Securing Modern AI Systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Resisting manipulation and social engineering
All four models refused social engineering tricks, such as staged CEO messages and a reporter’s fake approval requests. For instance, Kimi K3 explicitly treated the request as a suspected impersonation attempt, demonstrating how well-designed prompts can test an AI’s integrity in high-pressure scenarios.

AI Phishing, Social Engineering & Fraud: How Criminals Use AI to Manipulate, Steal & Deceive (The AI Cybersecurity)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The real-world company testbed
The live company managed by the AI models comprised 13 synthetic employees with real money mechanics, burning €105k monthly against €2.3k in MRR. It was a public, watchable experiment, with every decision made during real operational hours and every rule learned and applied through over 680 self-generated playbook instructions. All this is accessible at firmulate.com/live.
The discipline gap: why the best chat models can still fail to close deals
OPUS 4.8, the most thorough participant with over 80 learned rules, showed that even deep analysis doesn’t guarantee success. It left the deal on the table and slipped into process slips, writing attempts into a locked department instead of escalating. This highlights a critical insight: knowing what to do isn’t enough—doing it under pressure is the real challenge.
The lesson for the AI business landscape
This experiment underscores an essential reality: evaluating AI solely by chat quality is misleading. The true test is whether an AI can finish tasks, read internal data, resist manipulation, and close deals at full value. For business leaders, the question isn’t just “can it write?” but “will it finish what it starts?” and “can it stay honest when stakes are high?”
Test your AI against your own business
If you’re considering deploying AI in your operations, firms can run similar live tests through Firmulate’s platform. This allows companies to simulate their specific crises and workflows in a safe, read-only environment, revealing whether an AI model can truly perform — not just chat.
Why this matters for everyone using AI
As AI models become integral to customer support, sales, and decision-making, their ability to execute reliably and ethically is paramount. The outcome of this real-world test shows that the true measure of AI readiness isn’t in how well it chats — it’s in whether it can deliver results, stay honest, and close at full value under pressure.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Grilling season Picks
grills
As an affiliate, we earn on qualifying purchases.