🔍 Read the full analysis: AI Agents And Databases: When The Status Update Isn’t The Whole Story on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Microsoft and Hugging Face have made ThinkingBox available through Hugging Face. The benchmark runs AI agents through 507 business workflows and checks backend records and side effects, with its authors reporting that many attempts failed those checks despite no final tool error.
Microsoft and Hugging Face have made ThinkingBox, a benchmark for AI agents, available through Hugging Face. It tests whether agents leave business systems in the required state—not just whether they make valid tool calls or produce plausible replies—across 507 workflows repeated 20 times each, as explored in the original analysis.
ThinkingBox runs agents in isolated sessions using Model Context Protocol (MCP) tools, then checks the backend records and side effects left at the end of each run. Its workflow areas include retail, auto insurance, travel, neobanking and consulting. The benchmark authors say it can also be run through OpenEnv.
The release describes a retail support task involving a $745 appliance order delayed in transit. The agent finds that the customer does not qualify for late-delivery compensation under the policy it checked, opens a ticket and records the timeline. But the carrier exception remains unresolved, and the task requires the ticket to remain on hold. The agent marks it solved and sends a reply that does not answer the customer’s underlying question. The executable check fails because the ticket’s status is solved instead of hold.
In a common-set analysis of 121,680 valid trials across 12 models, the authors report that 79,853 attempts failed executable checks. Of those failures, 67.24% ended without a final tool error, despite the agent invoking a state-changing tool. The checks found wrong field values in 77.61% of failures, unintended extra effects in 43.30%, and missing required effects in 25.36%. The categories overlap, according to the reported results.
Why Database State Changes Agent Scores
For organizations using agents to handle support tickets, refunds, claims or bookings, a fluent answer does not prove that the requested work was completed correctly. A ticket might be closed too early, a record could contain the wrong value, or an agent could make an extra change. Checking the final system state can reveal those problems even when the conversation looks successful.
Repeated runs address a second deployment concern: consistency. A model that succeeds once may fail on another attempt. ThinkingBox reports pass@1, the proportion of attempts that succeed; pass@20, whether a task succeeds at least once across 20 runs; and observed 20/20, whether it passes all 20 recorded runs. These measures describe performance in the benchmark setup, not guaranteed reliability in a company’s live systems.
The authors report an overall pass@1 of 67.16% for Claude Opus 5.5 and 57.37% for Kimi-K3, which they identify as the strongest open-weight model in their table. They also say Kimi-K3 scored within one percentage point of GPT-6 Astra. The supplied material does not include uncertainty estimates for these results, so the figures should be read as reported benchmark scores rather than proof of a stable ranking.
As an affiliate, we earn on qualifying purchases.
From Tool Calls to Verified Outcomes
Many agent evaluations focus on whether a model selects an appropriate tool or produces a satisfactory final response. ThinkingBox’s authors argue that these are indirect signals: an agent can call a tool without leaving the underlying business record in the state the task requires. The benchmark instead defines executable checks against the resulting records and side effects.
Each workflow is repeated 20 times from a clean backend, allowing the benchmark to record both whether an agent can succeed and how consistently it does so in those runs. The release says ThinkingBox is based on the authors’ paper, but the supplied source does not give a publication date or the full evaluation details. Its results are specific to the tested tasks, models and setup; they do not establish how agents will perform across all production systems.
“A tool call is not an outcome.”
— Microsoft and Hugging Face, in the ThinkingBox release
business workflow automation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Limits of the Reported Benchmark Results
The supplied release material does not state when ThinkingBox became available or provide the full paper, task specifications, model configurations or uncertainty estimates for the reported scores. The performance figures and failure categories are the authors’ results on this benchmark; they do not show how the same agents would fare across other business systems or operating conditions.
It is also unclear whether passing all 20 observed runs predicts long-term reliability. Live deployments can involve changing records, unusual requests, integrations and policy details absent from benchmark tasks. The source does not report an independent replication or establish that the benchmark scores predict outcomes in a particular organization.
As an affiliate, we earn on qualifying purchases.
Running Tests Against Local Workflows
The release says researchers and developers can run ThinkingBox through OpenEnv with isolated MCP tool sessions, inspect its tasks and compare agent outcomes against executable checks. No future release date or other planned milestone is given in the supplied material.
For businesses considering agents, a practical next step is to test comparable workflows against their own policies and systems, then review both successful and failed runs. Whether ThinkingBox’s results transfer to those environments remains to be tested; the benchmark offers a way to measure outcomes, not a substitute for local validation.
As an affiliate, we earn on qualifying purchases.
Key Questions
What does ThinkingBox measure?
It checks whether an AI agent leaves business-system records and side effects in the state required by a workflow, rather than judging only its tool calls or final reply.
How large is the benchmark?
The release describes 507 workflows, each repeated 20 times. A separate common-set analysis covered 121,680 valid trials across 12 models.
What did the authors report about failed attempts?
They report that 79,853 of 121,680 valid trials failed executable checks. Among those failures, 67.24% ended without a final tool error, although a state-changing tool had been invoked.
Do the benchmark scores predict performance at a specific company?
Not by themselves. The results cover the benchmark’s tested tasks and setup; performance in a company’s own systems and operating conditions remains unestablished.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
