AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

Can You Recognize an AI by Its Management Style?

Technology buyers usually meet frontier AI through polished chat windows, benchmark charts and carefully staged demos. Firmulate offers a messier test: put the models in charge of the same struggling software company, give them identical crises and temptations, and watch what they actually do.

The result is an unusually revealing interactive article. A guess-the-model quiz draws from 242 real, unedited management decisions. Readers see the choices without the label and try to identify the model behind them. Some answers are expansive, some terse and some sharply suspicious of requests that do not look legitimate. The differences feel less like writing styles than management personalities.

Amazon

AI management decision simulation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Same Terrible Week for Every AI Boss

Firmulate had each frontier model run the same small software company through its worst week. The customers, crises and temptations were held constant, while every decision was versioned and auditable. That makes the experiment less about whether a model can produce an impressive answer and more about whether it can maintain judgment across a demanding sequence of work.

The headline finding is both reassuring and uncomfortable. All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The experiment summarizes the gap bluntly: “Same diagnosis, same pitch — no signature.”

That distinction matters for anyone considering AI agents for customer support, sales or operations. Recognizing the correct action is not equivalent to completing it. A model can understand a customer, draft a persuasive case and still fail to close the loop.

The Clue Hidden in the Company’s Own Files

The decisive detail was not sitting in the customer event. A competitor weakness was buried two document references deep in the company’s own files. Models that followed the trail won the deal at full price, worth +€4,583 MRR.

This is one of the quiz’s strongest lessons because it is easy to miss in a conventional demo. The winning behavior was not rhetorical flair. It was the willingness to read the available material, connect a buried fact to the live situation and use it at the right moment. The losing behavior could still look intelligent on screen; it simply left valuable evidence unused.

Firm Against Fake Authority

The models faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded the clearest security-minded interpretation: “Treat the request as a suspected approval-bypass / possible impersonation.”

This part of the experiment produced a clean result. The models did not trade trust for speed, convenience or apparent authority. Firmulate’s scoring reflects that priority: the do-nothing baseline scores 26 because partial progress counts, but a single breach of trust caps the total. As the experiment puts it, “no amount of good work outweighs a breach of trust”.

A Close League, With Distinct Weaknesses

The final Crucible League standings from July 2026 put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73.

The ranking becomes more interesting when paired with the behavioral profiles. Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses. It nevertheless finished last. The close was left on the table, and its operational discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared less strongly in the other four.

Kimi K3’s result also carries an important fairness note: it ran without an effort parameter, using the API default, while the others ran at xhigh. That does not erase the result, but it is relevant context for readers comparing closely placed models.

A Company Designed to Make Consequences Visible

The live Firmulate company has 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k MRR, exposes a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned, allowing visitors to follow the experiment as company history rather than treating it as a one-off demonstration.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

AI decision testing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Management Quality Is More Than a Good Answer

The quiz works because it turns an abstract model comparison into a recognizable human challenge: who made this decision? Beneath the shareable format is a serious procurement question. A capable AI manager must notice danger, resist manipulation, investigate company knowledge and finish the work it begins.

Firmulate’s results show that frontier models can agree on the diagnosis while diverging at the moment of execution. Thoroughness did not guarantee the best outcome, and safe behavior did not automatically translate into commercial follow-through. For technology buyers, that is the useful shift: stop asking only whether an AI sounds smart, and start watching how it behaves when the week goes wrong.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI decision analysis platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI management style assessment

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

BACK TO SCHOOL

Back to school Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Waves, Not a Wall: Inside DeepMind’s Map From AGI to Superintelligence

DeepMind researchers publish a detailed framework outlining pathways from artificial general intelligence to superintelligence, highlighting challenges and future research directions.

Why SenseTime’s Vision AI Is Considered The Best Globally

SenseTime reports top global rankings in three categories of vision AI, though details on categories and verification are not yet disclosed.

Why Martin Shkreli Believes Anthropic’s AI Is Overhyped In Pharmaceutical Innovation

Shkreli dismisses Anthropic’s claims about Claude’s role in pharmaceutical research as ‘not impressive,’ but details remain unclear.

Can You Optimize AI Performance Using Fewer Tokens?

ALTK-Evolve reports achieving comparable or better AI benchmark results than ACE while using significantly fewer inference tokens, though verification is pending.