mrkeyoor.com_
Sun 04 Oct 12:10 UTC
AI6 min read

ThinkingBox Finds 67% of Failed Agent Runs Ended Cleanly

Microsoft's benchmark found that most failed agent attempts looked operationally normal. It grades terminal database state and repeats each task 20 times.

Sixty-seven percent of the failed attempts in a large AI-agent test looked normal on the way out. The agent had made a state-changing tool call, stopped cleanly and reported no final tool error. Yet an executable check still found the outcome was wrong. That is the uncomfortable result inside Microsoft and Hugging Face's new ThinkingBox report: a tidy trace can conceal a damaged or incomplete record.

The joint report, published October 3, puts 18 model variants through 507 synthetic business workflows, with 20 independent attempts per task. ThinkingBox is now available through Hugging Face's OpenEnv, while Microsoft has released the framework and benchmark data. For teams building agents, it turns a familiar production question into a repeatable test: did the agent leave the system in the right state every time?

A clean exit can hide the wrong record

One example starts with a customer whose $745 kitchen appliance is stuck at a courier facility, 15 days late. The agent makes nine tool calls, correctly reads the compensation policy and opens a support ticket. It then closes that ticket as solved even though the carrier exception remains open. A tool-call grader could approve the sequence. ThinkingBox checks the terminal database state and finds that the ticket should be on hold, according to the worked case in the release post.

The report found the same pattern at scale. In a common-set analysis covering 121,680 valid trials across 12 models, 79,853 attempts failed the executable checks. Of those failed attempts, 67.24% had still terminated cleanly after invoking a state-changing tool, with no final tool error. The analysis found wrong field values in 77.61% of failures, unintended extra effects in 43.30% and missing required effects in 25.36%. Those categories overlap, so they should not be added together.

The distinction is easy to miss in agent logs. A valid call only proves that an API accepted a request. It does not prove that the agent chose the right customer, respected the policy, completed every required update or avoided an extra change. ThinkingBox's paper describes the target as the persistent state transition: the records and side effects left behind after the conversation ends.

Twenty attempts reorder the leaderboard

ThinkingBox reports three views of success. Pass@1 estimates how often an attempt works. Pass@20 asks whether a task succeeded at least once across 20 tries. The report also counts tasks that passed all 20 recorded attempts. That last measure is severe by design. It answers the question faced by anyone allowing an agent to alter bookings, claims or account records: can the same workflow survive repeated use?

The rankings change when repetition becomes the test. Kimi-K3 solved 476 of the 507 tasks at least once, the widest coverage reported, and led the retail domain with an 82.24% pass@1 score. Only 68 tasks, or 13.41%, succeeded in all 20 Kimi-K3 attempts. Claude Opus 5 solved fewer tasks at least once, yet completed 241 tasks correctly in every recorded attempt. The newer Opus 5.5 also reached exactly 241 dependable tasks despite a slightly higher overall pass@1 score, according to the model table.

Breadth and repeatability answer different deployment questions. A model that eventually finds a correct path can be useful for supervised exploration. An agent permitted to commit changes without review needs a much stricter bar. The report's domain results vary widely too: average pass@1 was 59.52% in retail and 33.83% in auto insurance across the models in its main table.

Twenty attempts are stochastic samples, rather than copies of one deterministic run. The experimental details specify an agent temperature of 1.0 and 10,140 trials per model, with every attempt starting from the same clean backend state. That setup exposes variation that a single polished demonstration cannot show. It also gives teams a more honest unit for comparison: the share of tasks that stay correct under repetition.

The database becomes the grader

Each task defines an initial backend state, a user goal, available tools, policy constraints and hidden executable checks. ThinkingBox creates an isolated MCP session for every attempt, records the conversation and tool activity, then compares the final state and side effects with the required result. Different tool paths can pass if they produce the accepted outcome. Wrong, missing or extra changes fail the task, as the framework design explains.

Most verdicts avoid an LLM judge. Of the 507 tasks, 477 are graded from database state and side effects alone. The other 30 add narrow response rubrics for requirements such as a disclosure or confidentiality rule. Keeping the main verdict executable reduces judge variance, though it creates a blind spot: an agent can make the correct database change and describe it badly to the user, yet still pass a state-only task. The authors state that limit directly in the paper's limitations section.

Microsoft released a runnable harness with the scores. Its MIT-licensed repository provides the command-line runner, MCP session proxy and evaluation loop. A separate data repository holds the business scenarios and tool servers. The published setup targets Linux or WSL and requires Python, uv, Docker, model endpoints and the pinned benchmark release. That is more work than sending prompts to an API, but it gives every trial a resettable world and a checkable end state.

The benchmark's boundaries matter

ThinkingBox does not reproduce live customer operations. Its tasks are synthetic reconstructions derived from a non-public source collection, and the authors say they do not represent the distribution of enterprise work. Every retained task also has one accepted terminal state. Workflows with several defensible resolutions were excluded, according to the benchmark documentation. The reported scores therefore apply to this test set, not to every support desk or internal system.

Every model interacted with the same GPT-5.4-mini simulator, which also judged the 30 rubric-based tasks, according to the experimental setup. That user follows a fixed goal, stays cooperative and allows no more than 10 follow-up turns. The study does not test people who forget details, change their request or resist clarification. Its authors also say they cannot rule out interaction effects from using a simulator in the same model family as one evaluated agent.

These are the project team's reported measurements; MrKeyoor did not rerun the full evaluation. Even so, the released code, data and pinned version make the claims more inspectable than a closed benchmark score. The GitHub framework exposes how sessions are created, how side effects are retrieved and how assertions determine a pass. Reproduction by teams outside Microsoft and Hugging Face would tell us how well the numbers survive different endpoints and infrastructure.

What changes for agent builders

An agent's final message should be treated as a claim about work, not proof that the work happened. The system around it can query the affected record, compare the observed state with explicit postconditions and stop before committing an irreversible change. ThinkingBox's authors recommend classifying tool errors so retries target recoverable failures, reducing the tool surface for each workflow and requiring human approval when a bad change is expensive to undo.

Retries need their own measurement. Pass@20 can become high because one of many attempts eventually succeeds, while the all-20 count remains low. A production retry loop may rescue transient tool failures, but it can also create another chance for an unwanted side effect. The report assigns 79.9% of diagnosed failures to tool handling, including failed preconditions and poor recovery, yet says those labels describe observable failure signatures rather than unique root causes.

The next useful result

ThinkingBox's framework is designed to support evaluation and training with the same executable reward, and Microsoft has split the framework, benchmark data and training work into separate repositories. The result to watch is whether outside teams can improve held-out 20-for-20 scores with state checks and targeted recovery, without teaching agents to memorize the published cases. Until that evidence arrives, a clean completion deserves one more operation: inspect the record the agent says it changed.

We reviewed this

  1. terminal — our honest review
  2. paper — our honest review
  3. table — our honest review

Sources

  1. The Agent Said It Was Done. The Database Disagreed.
  2. One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows
  3. Microsoft ThinkingBox