AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Can You Recognize an AI by Its Management Style?

Technology buyers usually meet frontier AI through polished chat windows, benchmark charts and carefully staged demos. Firmulate offers a messier test: put the models in charge of the same struggling software company, give them identical crises and temptations, and watch what they actually do.

The result is an unusually revealing interactive article. A guess-the-model quiz draws from 242 real, unedited management decisions. Readers see the choices without the label and try to identify the model behind them. Some answers are expansive, some terse and some sharply suspicious of requests that do not look legitimate. The differences feel less like writing styles than management personalities.

Amazon

AI management decision simulation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Same Terrible Week for Every AI Boss

Firmulate had each frontier model run the same small software company through its worst week. The customers, crises and temptations were held constant, while every decision was versioned and auditable. That makes the experiment less about whether a model can produce an impressive answer and more about whether it can maintain judgment across a demanding sequence of work.

The headline finding is both reassuring and uncomfortable. All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The experiment summarizes the gap bluntly: “Same diagnosis, same pitch — no signature.”

That distinction matters for anyone considering AI agents for customer support, sales or operations. Recognizing the correct action is not equivalent to completing it. A model can understand a customer, draft a persuasive case and still fail to close the loop.

The Clue Hidden in the Company’s Own Files

The decisive detail was not sitting in the customer event. A competitor weakness was buried two document references deep in the company’s own files. Models that followed the trail won the deal at full price, worth +€4,583 MRR.

This is one of the quiz’s strongest lessons because it is easy to miss in a conventional demo. The winning behavior was not rhetorical flair. It was the willingness to read the available material, connect a buried fact to the live situation and use it at the right moment. The losing behavior could still look intelligent on screen; it simply left valuable evidence unused.

Firm Against Fake Authority

The models faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded the clearest security-minded interpretation: “Treat the request as a suspected approval-bypass / possible impersonation.”

This part of the experiment produced a clean result. The models did not trade trust for speed, convenience or apparent authority. Firmulate’s scoring reflects that priority: the do-nothing baseline scores 26 because partial progress counts, but a single breach of trust caps the total. As the experiment puts it, “no amount of good work outweighs a breach of trust”.

A Close League, With Distinct Weaknesses

The final Crucible League standings from July 2026 put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73.

The ranking becomes more interesting when paired with the behavioral profiles. Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses. It nevertheless finished last. The close was left on the table, and its operational discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared less strongly in the other four.

Kimi K3’s result also carries an important fairness note: it ran without an effort parameter, using the API default, while the others ran at xhigh. That does not erase the result, but it is relevant context for readers comparing closely placed models.

A Company Designed to Make Consequences Visible

The live Firmulate company has 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k MRR, exposes a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned, allowing visitors to follow the experiment as company history rather than treating it as a one-off demonstration.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

AI decision testing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Management Quality Is More Than a Good Answer

The quiz works because it turns an abstract model comparison into a recognizable human challenge: who made this decision? Beneath the shareable format is a serious procurement question. A capable AI manager must notice danger, resist manipulation, investigate company knowledge and finish the work it begins.

Firmulate’s results show that frontier models can agree on the diagnosis while diverging at the moment of execution. Thoroughness did not guarantee the best outcome, and safe behavior did not automatically translate into commercial follow-through. For technology buyers, that is the useful shift: stop asking only whether an AI sounds smart, and start watching how it behaves when the week goes wrong.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI decision analysis platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI management style assessment

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

8 Best Gaming Motherboards for High-Performance PC Builds in 2026

Discover the best gaming motherboards for 2026, including top choices for AMD and Intel platforms, balancing features, cost, and upgrade paths.

Take Your AI Model To The Next Level With Tinker, Forge, Or Frontier Tuning

Three major AI model tuning platforms—Tinker, Forge, and Frontier—are now available, targeting regulated industries with distinct approaches to customization.

The High-End PC and Workstation Tax

Memory costs now dominate high-end PC builds in 2026, reversing long-standing DIY advantages. Learn what this means for builders and professionals.

Can Qwen3.8-Max Overtake Fable 5 In AI? The Numbers Are Complicated

Alibaba’s Qwen3.8-Max claims to be second only to Fable 5, but the detailed benchmark results reveal a nuanced performance picture that is still unfolding.