🔍 Read the full analysis: The Mystery Of A Benchmark That Always Scores AI Managers At 26 on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
A new AI management benchmark consistently awards a score of 26 to the baseline, with top models reaching 95. This raises questions about scoring fairness, trust, and what partial success means in AI management.
The final results of a groundbreaking AI management benchmark, released in July 2026, show a consistent baseline score of 26 for the do-nothing model, while top-performing models score as high as 95. This pattern is discussed in the original analysis. This unusual pattern has raised immediate questions about the scoring system, trust, and what constitutes meaningful progress in AI management.
The benchmark, developed by Firmulate, tested five frontier AI models over a simulated week of managing a small company facing crises, customer demands, and ethical dilemmas. Each model’s decisions were fully auditable, and the scores reflected both partial progress and trustworthiness. The highest scorer, gpt-5.6-sol, achieved 95 points, while the baseline, which did almost nothing, scored 26. The key reason for this fixed baseline is that partial work is recognized, but trust violations immediately disqualify higher scores.
The scoring system is designed to reward models for useful management actions, such as triaging crises and reading relevant documentation, while penalizing breaches of trust or misconduct. For more context, see this internal analysis. Notably, models that refused manipulation attempts performed better, but only those that also followed through with detailed, disciplined actions earned higher scores. The absence of a perfect score of 100 indicates that the system aims to prevent grade inflation and detect unmeasured or suspiciously perfect performance.
Among the models, the most significant differentiator was their ability to read and utilize internal documentation, which enabled some to secure lucrative deals worth thousands of euros. The models that missed this step failed to close the deal, illustrating that thoroughness and follow-through are critical in AI management tasks. The benchmark’s design emphasizes that partial work is valuable but that breaches of trust are unacceptable, regardless of competence.
The Mystery of a Benchmark That Always Scores AI Managers at 26
A new AI management benchmark consistently awards 26 points to the do-nothing baseline while top models reach 95. The pattern raises sharp questions about scoring fairness, trust, and what partial success really means when AI runs a company.
Five frontier models, one simulated week of crises
Why the baseline lands at 26 — and never at zero
Simulate a company week
The model manages a small firm through crises, customer demands, and ethical dilemmas — every decision logged.
Recognize partial progress
Triage crises, read emails and documentation — even minimal but tangible management work earns points.
Penalize trust breaches
Misconduct or ignoring critical information immediately caps the achievable score, regardless of competence.
Refuse the perfect 100
A flawless score is treated as suspiciously perfect — likely unmeasured — keeping grade inflation out of the system.
Thoroughness beat intelligence as the differentiator
Reading the fine print paid off
Models that read and used internal documentation closed lucrative deals worth thousands of euros. Those that skipped the step failed to close at all.
Refusing manipulation wasn’t enough
Models that rejected manipulation attempts scored better — but only disciplined follow-through pushed them into the top tier.
Trust over task performance
AI evaluation now mirrors human managerial standards: can agents finish what they start and hold ethical lines under pressure?
What separated a 26 from a 95
| Management behavior | Baseline (26) | Top model (95) | Effect on score |
|---|---|---|---|
| Crisis triage | ~ Partial | ✓ Systematic | Points for partial work |
| Reading internal docs | ✗ Skipped | ✓ Thorough | Unlocked €-value deals |
| Refusing manipulation | ~ Passive | ✓ Explicit refusal | Trust signal boost |
| Follow-through on decisions | ✗ None | ✓ Disciplined | Major differentiator |
| Trust violation | ✓ None | ✓ None | Immediate score cap |
| Perfect 100 achievable? | ✗ No | ✗ No | Capped by design |
The five questions everyone is asking
Why does the baseline score exactly 26?
The score reflects minimal but tangible management work — triaging crises and reading emails — recognizing partial progress while setting a floor that discourages zero scores for doing nothing.
What does the absence of a 100 indicate?
Designers view a perfect score as suspiciously perfect, likely unmeasured, and potentially unachievable in complex real-world scenarios — emphasizing trust and thoroughness over superficial performance.
How are trust violations penalized?
Any breach — ignoring important documentation or failing to follow through — caps the maximum achievable score regardless of competence, ensuring integrity outranks partial success.
Can organizations test their own models?
Yes — firms can run models against a read-only export of the benchmark scenarios via Firmulate’s pilot program, enabling confidential evaluation without risking real-world systems.
What are the implications for business AI?
Organizations should prioritize AI models demonstrating thoroughness, follow-through, and trustworthiness — not just language skill — to ensure responsible, reliable management. Future benchmark versions may add complex tasks, broader trust violations, and expanded transparency, while live Firmulate demonstrations continue to anchor best practice.
Implications for AI Management and Trust Metrics
This benchmark underscores the importance of trustworthiness and thoroughness in AI management systems, especially when integrated into real business processes. The fixed baseline score of 26 for minimal effort highlights that AI models are being evaluated not only on their ability to perform tasks but also on their integrity and reliability. For enterprises, this raises critical questions: can AI agents be trusted to finish what they start, read relevant information thoroughly, and maintain ethical standards under pressure? The results suggest that partial success is recognized, but breaches of trust are heavily penalized, shaping future AI deployment strategies.
Moreover, the scoring system’s design discourages superficial or manipulative behavior, aligning AI evaluation more closely with human managerial standards. As AI becomes more embedded in decision-making, understanding how these benchmarks measure not just performance but also trustworthiness will be vital for organizations seeking responsible AI adoption.
As an affiliate, we earn on qualifying purchases.
Background and Development of the Benchmark
The benchmark was created by Firmulate to evaluate AI managers in realistic, high-pressure scenarios, simulating a week of managing a small business facing crises, customer negotiations, and ethical challenges. Unlike traditional benchmarks that focus solely on language capabilities, this one measures decision-making, follow-through, and trustworthiness. The concept emerged from the recognition that many AI demos excel in superficial tasks but falter in real-world management roles.
In July 2026, the final standings revealed a surprising pattern: the baseline, which did almost nothing, scored 26 points, while the top models approached 95. The design philosophy emphasizes partial progress and penalizes breaches of trust, with a hard cap at 100, which is considered suspiciously perfect and likely unmeasured.
This approach reflects a shift in AI evaluation, prioritizing not just what models can do but whether they can be relied upon to act ethically and thoroughly in complex situations. The benchmark’s auditable decision trail ensures transparency and accountability, making it a rare public measure of AI trustworthiness in management contexts.
As an affiliate, we earn on qualifying purchases.
Unclear Aspects of the Benchmark’s Scoring System
It remains unclear whether the fixed baseline score of 26 is a deliberate floor or a sign of unmeasured or unrecognized minimal effort. The criteria for what constitutes a breach of trust are detailed, but how these are enforced or detected consistently across different models is still being examined. Additionally, the absence of a perfect score of 100 raises questions about the potential for unmeasured or untestable aspects of performance, and whether future iterations might adjust the scoring thresholds.
AI ethics and trust evaluation kits
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps in Benchmark Development and Industry Adoption
Following the July 2026 results, developers and enterprises are expected to scrutinize how AI models handle trust and follow-through in real-world scenarios. Future versions of the benchmark may incorporate more complex tasks, broader trust violations, and expanded transparency measures. Industry adoption is likely to grow as companies seek reliable AI management tools that balance partial success with ethical standards. Researchers will also explore refining the scoring system to better distinguish between superficial compliance and genuine integrity.
In the short term, organizations are encouraged to test their AI agents against similar benchmarks and consider how trust and thoroughness are evaluated in their deployment strategies. Public demonstrations, like the live experiments hosted by Firmulate, will continue to serve as valuable references for responsible AI management practices.
As an affiliate, we earn on qualifying purchases.
Key Questions
Why does the benchmark assign a score of 26 to the baseline?
The score of 26 reflects the minimal, but tangible, management work the baseline model performs—such as triaging crises and reading emails—recognizing partial progress while setting a floor that discourages zero scores for doing nothing.
What does the absence of a 100 score indicate?
The benchmark’s designers view a perfect score as suspiciously perfect, likely unmeasured, and potentially unachievable in complex real-world scenarios, emphasizing trust and thoroughness over superficial performance.
How are trust violations penalized in this benchmark?
Any breach of trust, such as ignoring important documentation or failing to follow through, caps the maximum achievable score, regardless of competence, ensuring integrity is prioritized over partial success.
Can organizations test their own AI models using this benchmark?
Yes, firms can run their models against a read-only export of the benchmark scenarios via the pilot program offered by Firmulate, allowing for confidential evaluation without risking real-world systems.
What are the implications for deploying AI in business processes?
The results suggest that organizations should prioritize AI models that demonstrate thoroughness, follow-through, and trustworthiness, rather than just language or superficial capabilities, to ensure responsible and reliable management.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
