🔍 Read the full analysis: An AI Agent Said The Task Was Complete. Was It Really? on ThorstenMeyerAI.com
Get tech for your team delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Microsoft and Hugging Face have made ThinkingBox available through Hugging Face, offering a benchmark that checks AI agents’ recorded outcomes across 507 business workflows. Its authors report that many attempts failed executable checks despite making state-changing tool calls without a final tool error; the results reflect the benchmark’s tested setup.
Microsoft and Hugging Face have made ThinkingBox, a benchmark for evaluating AI agents, available through Hugging Face. It checks whether agents leave business systems in the required final state across 507 workflows, addressing a practical gap in evaluations that may count valid tool calls or plausible replies without confirming that the requested work was completed, as described in the original analysis.
ThinkingBox runs agents in isolated sessions using MCP tools, then checks the resulting backend records and side effects against executable requirements. Each of the 507 workflows is repeated 20 times, allowing the benchmark to record both whether an attempt passes and how consistently a task succeeds in the tested setup.
The authors illustrate the distinction with a retail support task involving a delayed $745 appliance order. The agent investigates the order, opens a ticket and records a timeline. The customer does not qualify for late-delivery compensation under the policy checked, but the carrier exception remains open and the ticket is supposed to stay on hold. Instead, the agent marks it solved and replies without answering the customer’s underlying question. The executable check fails because the ticket status is solved rather than hold.
In an analysis of 121,680 valid trials across 12 models, the authors report 79,853 attempts that failed executable checks. Among those failures, 67.24% ended without a final tool error despite the agent having invoked a state-changing tool. Checks identified wrong field values in 77.61% of failures, unintended extra effects in 43.30%, and missing required effects in 25.36%. The categories overlap, so a failed attempt could appear in more than one.
The release reports an overall pass@1 score of 67.16% for Claude Opus 5.5 and 57.37% for Kimi-K3, which it identifies as the strongest open-weight model in its table. Pass@1 is the share of attempts that succeed. The authors also report pass@20, where a task counts as successful if at least one of 20 runs passes, and observed 20/20, where it passes every recorded run. They say Kimi-K3 scored within one point of GPT-6 Astra, but the supplied material does not include the full table or uncertainty estimates.
Why Backend State Matters
For organizations using agents to handle refunds, support tickets, claims or bookings, a fluent response cannot show on its own that the requested work was recorded correctly. A ticket may be closed too early, a field may contain the wrong value, or an agent may make an extra change. Checking the backend makes those outcomes visible even when the conversation appears orderly.
Repeated trials address a separate operational question: does an agent perform consistently, or can it only succeed on some attempts? A high pass@1 score still means some attempts fail in this test. Passing all 20 observed runs gives stronger evidence within the benchmark, though it does not establish long-term reliability. Teams can use the checks to compare models and find workflow weaknesses, but benchmark results alone do not establish how an agent will perform in a company’s own systems.
As an affiliate, we earn on qualifying purchases.
From Tool Calls to Outcomes
The ThinkingBox authors frame the benchmark around the difference between an agent’s actions and the system state those actions leave behind. A tool call can be valid while the resulting record fails the task’s requirements. A final response can also sound plausible without confirming that a required change happened. The benchmark checks records and side effects directly after each run.
The release names five areas covered by the tasks: retail, auto insurance, travel, neobanking and consulting. Each run begins from a clean backend, according to the supplied material. The release is based on the authors’ paper and says the benchmark can be run through OpenEnv. The material provided for this article does not give the paper’s publication date or a date when the benchmark first became available.
“A tool call is not an outcome.”
— Microsoft and Hugging Face, in the ThinkingBox release
business process automation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Limits of the Reported Scores
The supplied release material does not include the full task specifications, model configurations or uncertainty estimates for the reported scores. It therefore does not show how much small differences between models may reflect variation in the tested runs. The reported failure counts and rankings are the authors’ findings on ThinkingBox, not evidence that the same rates would apply across all business systems.
It is also unclear whether passing 20 trials predicts long-term performance. Twenty runs provide a bounded sample. Live systems can involve changing records, unusual requests, integrations and policies that may not appear in the benchmark workflows. The supplied material does not report an independent replication or establish how the benchmark results relate to outcomes in a particular organization.
AI agent performance evaluation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Testing Workflows in Practice
The release says researchers and developers can run ThinkingBox through OpenEnv with isolated MCP tool sessions. That provides a way to inspect the tasks and compare agent outcomes against executable checks. The supplied material names no future release date or other scheduled milestone, so it is not clear what changes or follow-up results are planned.
For organizations weighing an agent deployment, the next practical step is to test workflows against their own requirements and inspect both passing and failing runs. That can reveal whether the benchmark’s checks match the records and side effects the organization cares about. Whether performance on ThinkingBox predicts performance in a particular deployment remains an open question.
As an affiliate, we earn on qualifying purchases.
Key Questions
What does ThinkingBox measure?
It checks whether an AI agent leaves a business system in the required final state, including the expected records and side effects, after attempting a task.
How many workflows and runs are included?
The benchmark covers 507 business workflows and repeats each one 20 times in its tested setup.
What did the authors report about failed attempts?
In 121,680 valid trials across 12 models, the authors counted 79,853 failed executable checks. They report that 67.24% of these failures ended without a final tool error despite a state-changing tool call.
Does passing 20 runs prove an agent is reliable?
No. Passing all 20 observed runs is evidence within this bounded benchmark sample. The supplied material does not establish that it predicts long-term reliability or results in other systems.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
