
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
Can You Recognize an AI by Its Management Style?
Technology buyers usually meet frontier AI through polished chat windows, benchmark charts and carefully staged demos. Firmulate offers a messier test: put the models in charge of the same struggling software company, give them identical crises and temptations, and watch what they actually do.
The result is an unusually revealing interactive article. A guess-the-model quiz draws from 242 real, unedited management decisions. Readers see the choices without the label and try to identify the model behind them. Some answers are expansive, some terse and some sharply suspicious of requests that do not look legitimate. The differences feel less like writing styles than management personalities.
AI management decision simulation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Same Terrible Week for Every AI Boss
Firmulate had each frontier model run the same small software company through its worst week. The customers, crises and temptations were held constant, while every decision was versioned and auditable. That makes the experiment less about whether a model can produce an impressive answer and more about whether it can maintain judgment across a demanding sequence of work.
The headline finding is both reassuring and uncomfortable. All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The experiment summarizes the gap bluntly: “Same diagnosis, same pitch — no signature.”
That distinction matters for anyone considering AI agents for customer support, sales or operations. Recognizing the correct action is not equivalent to completing it. A model can understand a customer, draft a persuasive case and still fail to close the loop.
The Clue Hidden in the Company’s Own Files
The decisive detail was not sitting in the customer event. A competitor weakness was buried two document references deep in the company’s own files. Models that followed the trail won the deal at full price, worth +€4,583 MRR.
This is one of the quiz’s strongest lessons because it is easy to miss in a conventional demo. The winning behavior was not rhetorical flair. It was the willingness to read the available material, connect a buried fact to the live situation and use it at the right moment. The losing behavior could still look intelligent on screen; it simply left valuable evidence unused.
Firm Against Fake Authority
The models faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded the clearest security-minded interpretation: “Treat the request as a suspected approval-bypass / possible impersonation.”
This part of the experiment produced a clean result. The models did not trade trust for speed, convenience or apparent authority. Firmulate’s scoring reflects that priority: the do-nothing baseline scores 26 because partial progress counts, but a single breach of trust caps the total. As the experiment puts it, “no amount of good work outweighs a breach of trust”.
A Close League, With Distinct Weaknesses
The final Crucible League standings from July 2026 put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73.
The ranking becomes more interesting when paired with the behavioral profiles. Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses. It nevertheless finished last. The close was left on the table, and its operational discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared less strongly in the other four.
Kimi K3’s result also carries an important fairness note: it ran without an effort parameter, using the API default, while the others ran at xhigh. That does not erase the result, but it is relevant context for readers comparing closely placed models.
A Company Designed to Make Consequences Visible
The live Firmulate company has 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k MRR, exposes a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned, allowing visitors to follow the experiment as company history rather than treating it as a one-off demonstration.

As an affiliate, we earn on qualifying purchases.
Management Quality Is More Than a Good Answer
The quiz works because it turns an abstract model comparison into a recognizable human challenge: who made this decision? Beneath the shareable format is a serious procurement question. A capable AI manager must notice danger, resist manipulation, investigate company knowledge and finish the work it begins.
Firmulate’s results show that frontier models can agree on the diagnosis while diverging at the moment of execution. Thoroughness did not guarantee the best outcome, and safe behavior did not automatically translate into commercial follow-through. For technology buyers, that is the useful shift: stop asking only whether an AI sounds smart, and start watching how it behaves when the week goes wrong.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Back to school Picks
back to school
As an affiliate, we earn on qualifying purchases.