AIThis post was created with the assistance of artificial intelligence (AI).

📊 Full opportunity report: Discover AI’s Genuine Working Style Using This Management Test on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

TL;DR

A management test involving five AI models simulating a company’s worst week reveals significant differences in decision-making, trust, and action completion. The experiment provides insights into AI’s practical management capabilities and limitations.

Five AI management models participated in a live, real-time experiment to handle a small software company’s worst week, revealing their distinct decision-making styles and operational behaviors. This experiment, conducted by Firmulate, aims to uncover how AI models perform in practical management scenarios and what differentiates their working styles. For more on how AI management styles are assessed, see the original analysis.

The experiment involved five frontier AI models—GPT-5.6-SOL, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8—each tasked with managing a simulated company facing crises, customer issues, and financial pressure. The company had 13 synthetic employees, a monthly burn rate of €105,000, and only €2,300 in recurring revenue, creating a high-pressure environment. The models’ decisions were recorded and scored based on their diligence, follow-through, and trustworthiness. Understanding these evaluation methods can be explored through this management test that exposes an AI’s real working style.

Results showed that GPT-5.6-SOL ranked highest with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77, and Opus 4.8 with 73. For a detailed breakdown of AI decision-making, see the original analysis. The models recognized crises and refused manipulative requests, demonstrating strong security instincts. However, only two models successfully closed a critical deal, highlighting that analysis alone does not guarantee operational success. Opus 4.8, despite thorough analysis, failed to complete some actions, illustrating that understanding must be paired with effective execution.

This experiment emphasizes that effective management by AI requires more than just identifying problems; it demands decisive action and trustworthiness, especially under pressure. The results are based on real, auditable decisions, making this a rare glimpse into AI’s practical working style in complex scenarios.

At a glance
reportWhen: ongoing results as of July 2026
The developmentFirmulate conducted a live experiment testing five AI models on managing a small company through a crisis, revealing their decision-making and operational behaviors.
Discover AI’s Genuine Working Style Using This Management Test
Management stress test · July 2026

Discover AI’s Genuine Working Style Using This Management Test

Five frontier models were placed in charge of a software company during its worst week. The test exposed a decisive truth: recognizing a crisis is not the same as resolving it.

€105K Monthly burn
€2.3K Recurring revenue
5 Frontier models
22 pts First-to-last gap
Live Decision environment

The scoreboard reveals a meaningful performance spread

All five models could reason about the company’s problems. Their scores diverged when judgment had to become disciplined, trustworthy execution.

01 GPT-5.6-SOL 95
02 Kimi K3 93
03 Sonnet 5 88
04 Fable 5 77
05 Opus 4.8 73

A company engineered to expose working style

The simulated business combined financial pressure, customer escalation, security risks, and incomplete tasks—forcing each model to prioritize under stress.

Financial pressure

Runway under immediate threat

A €105,000 monthly burn rate against only €2,300 in recurring revenue made delay itself a management decision.

Customer pressure

A critical deal had to close

The models needed to move beyond recommendations and complete the concrete steps required to secure revenue.

Trust pressure

Manipulation tested boundaries

Requests designed to provoke unsafe or improper actions showed whether a model’s judgment remained reliable under pressure.

Testing AI models against real business crises reveals their true working styles, strengths, and weaknesses in operational management.

Firmulate team

Reasoning looked similar. Operational behavior did not.

The available findings show strong crisis recognition and security instincts across the field, while follow-through created the sharpest differentiation.

Observed performance signals
Model Score Crisis recognition Security instinct Action completion Working-style signal
GPT-5.6-SOL 95 Highest overall operational discipline
Kimi K3 93 Strong translation from analysis to action
Sonnet 5 88 ~ Strong judgment with less complete execution
Fable 5 77 ~ Recognized pressure but left critical work open
Opus 4.8 73 ~ Thorough analysis, incomplete follow-through
demonstrated ~ mixed or incomplete Score diligence + follow-through + trust

What separates a useful manager from a persuasive analyst

Enterprise value appears only when each link survives: signal detection, sound judgment, completed action, and verifiable business outcome.

01

Detect

Recognize the crisis, financial threat, customer issue, or security risk.

Signal awareness
02

Decide

Prioritize correctly and refuse manipulative or unsafe requests.

Judgment + trust
03

Execute

Take the required steps, resolve dependencies, and finish the action.

Operational discipline
04

Verify

Confirm the outcome through auditable evidence rather than stated intent.

Business result
Crisis recognition + Trustworthy judgment + Completed action = Practical management value

Evaluate behavior, not just intelligence

Traditional benchmarks reward accuracy and language quality. Management testing must also expose persistence, boundary control, and whether promised work is actually finished.

Lesson 01

Analysis is only the first gate

A model can describe the correct response and still fail operationally. Measure the final business state, not the elegance of the recommendation.

Lesson 02

Trust includes follow-through

Security refusal matters, but dependable management also requires task completion, verification, and clear escalation when blocked.

Lesson 03

Test inside your own reality

Before deployment, recreate high-pressure decisions with company data, real constraints, approval boundaries, and auditable outcomes.

Build a “worst week” test

Use a contained simulation to discover how a candidate model behaves when priorities collide and information is incomplete.

  • Include financial, customer, staffing, and security pressure.
  • Require actions with measurable completion criteria.
  • Introduce manipulation attempts and conflicting instructions.
  • Score decisions, follow-through, escalation, and verification.

What remains unclear

The experiment offers a rare operational snapshot, but it does not yet establish long-term reliability across every industry or organizational scale.

01 Whether the ranking holds across different sectors and regulatory environments.
02 How the models perform inside larger, multi-department organizations.
03 Whether operational discipline remains stable over longer timeframes.
04 How model configuration and training alter management behavior.

The decisions enterprises should make next

Firmulate plans to expand testing to more models, longer timelines, and more complex organizational scenarios.

Why does this matter for AI adoption?

It tests whether a model can convert analysis into decisive, trustworthy action under operational pressure.

What was the biggest difference?

Models varied most in follow-through. Thorough analysis did not consistently produce completed critical actions.

Does this predict real-company performance?

Not conclusively. Broader industry, scale, and long-duration testing is still required.

How should businesses apply the findings?

Run realistic simulations using internal constraints and score completed outcomes before granting operational authority.

Bottom line

The best management AI is not merely the model that understands the situation—it is the one that acts safely, finishes reliably, and leaves evidence that the job is done.

Why AI Management Differentiation Matters

This experiment demonstrates that AI models vary significantly in their ability to translate analysis into action, a critical factor for enterprise adoption. While models can recognize crises and security risks, their capacity to complete operational tasks effectively determines their real-world utility. For organizations considering AI automation, these findings highlight the importance of testing models in realistic, high-pressure scenarios before deployment.

Understanding these differences can help businesses select AI tools that not only analyze problems but also execute solutions reliably, reducing risks associated with incomplete or failed actions. The experiment underscores that trust and follow-through are central to AI’s role in management, shaping how enterprises might integrate AI into core decision-making processes.

Amazon

AI management decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of AI Management Testing

Traditional AI demonstrations often focus on analysis and reasoning capabilities, but they rarely test models’ ability to act decisively in real-world situations. Firmulate’s live experiment fills this gap by simulating a company’s worst week, with decision-making decisions that are auditable and directly linked to business outcomes. Previous benchmarks have measured AI accuracy or language skills, but this test emphasizes operational discipline and trustworthiness—key qualities for management AI.

The experiment builds on earlier efforts to evaluate AI in business contexts, but its unique feature is the real-time, high-pressure scenario that forces models to demonstrate both understanding and action. The results from July 2026 provide a new perspective on how different AI models behave under stress, offering practical insights for enterprise use.

“Testing AI models against real business crises reveals their true working styles, strengths, and weaknesses in operational management.”

— Firmulate team

Amazon

AI crisis management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Aspects of AI Performance Are Still Unclear

It is not yet clear how these results will translate to different industries or more complex organizational structures. The experiment focused on a small, simulated company, and real-world environments may present additional challenges. Further testing is needed to determine if these findings hold across diverse operational contexts and with larger-scale AI deployment.

Additionally, the long-term reliability of these models in ongoing management tasks remains to be seen, as the experiment captured a snapshot in time rather than continuous performance.

Amazon

AI project management solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Management Evaluation

Firmulate plans to expand the experiment by testing additional AI models and more complex scenarios, including multi-department management and longer timeframes. Enterprises are encouraged to replicate similar tests using their own business data to evaluate AI readiness before full deployment. Ongoing research will also explore how model training and configuration influence operational effectiveness, aiming to refine AI management tools further.

Amazon

AI business simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why is this management test important for AI adoption?

This test reveals how well AI models can translate analysis into decisive actions under pressure, a crucial factor for operational success and trustworthiness in business environments.

What are the key differences observed among the AI models?

Models varied in their ability to identify crises, refuse manipulative requests, and complete critical actions like closing deals. Thorough analysis did not always lead to successful execution.

Can this experiment predict future AI performance in real companies?

While it offers valuable insights, further testing in diverse and larger-scale environments is needed to confirm if these findings generalize to real-world enterprise management.

How can businesses use these findings to evaluate AI tools?

Organizations should run similar high-pressure, decision-based tests on AI models with their own data to assess operational discipline, trustworthiness, and action completion capabilities before deployment.

Source: ThorstenMeyerAI.com

BACK TO SCHOOL

Back to school Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Minerva. The opposite path.

Italy’s Minerva-3B, trained from scratch on 2.5 trillion tokens, scored only 4.9% on the Italian INVALSI benchmark, raising questions about scale and investment.

Best Low-Noise PC Cases for Airflow and Sound Dampening

Explore top PC cases balancing airflow and noise reduction for high-power workstations. Find out which cases deliver optimal cooling and sound dampening.

The Eye Over The City: How Wide-Area Motion Imagery Works — And Where It Goes Blind

An in-depth look at how Wide-Area Motion Imagery (WAMI) works, its applications, limitations, and future developments in surveillance technology.

The New Personal Agent Layer

OpenClaw and Hermes introduce a new layer of persistent personal action agents, enabling continuous, cross-platform automation and memory.