Overview
ThinkingBox is an agent evaluation sandbox released jointly by Microsoft and Hugging Face on October 3, 2026, together with a benchmark named ThinkingBox-Bench. Its grading principle is blunt: ignore what the agent said at the end, ignore how many tool calls it made, and look at what is left in the backend database when the run is over.
The 507 tasks span five enterprise workflow domains: retail, auto insurance, travel, neobank, and consulting. Each task defines a starting backend state, a user goal, the available MCP tools, and a domain policy, then runs 20 independent attempts from an identical clean state. Every attempt gets an isolated MCP session, and no two attempts share a database row or cached tool state, which is what makes a 20-trial comparison meaningful. Afterwards a side-effect extractor derives what actually changed, and deterministic judges compare it against the required end state: any trajectory producing the right outcome passes, while wrong, missing, or extra effects fail. Of the 507 tasks, 477 are graded on state alone and 30 add a binary rubric for requirements with no clean database value, such as whether the agent disclosed that a result is not guaranteed.
The motivating example from the authors makes the point well: a support agent handles a late delivery, makes nine tool calls that all look correct, reads the refund policy correctly, and closes the ticket as resolved. Two things are still wrong: the courier exception remains open and the customer's actual question was never answered. A grader checking only tool calls or the final sentence would score it as a success. The database disagrees.
Both the sandbox and the dataset are on Hugging Face. The harness is MIT-licensed, the dataset uses CDLA-Permissive-2.0, the benchmark is served through the OpenEnv interface, and each finished episode returns a binary pass or fail reward.
Key Features
- Grades the terminal database state: Judges compare the state and side effects actually left in the backend, accepting any trajectory that produces the right outcome and rejecting wrong field values, extra effects, and missing effects
- 20 independent reruns per task: Every attempt gets an isolated MCP session with freshly initialized state, and no two attempts share a database row or cached tool state, which is what makes comparing 20 trials meaningful
- Simulated user holds private context: A booking reference, a preference, or a date of birth is held by the simulated user and released only when asked, forcing the agent to elicit missing information
- Five enterprise workflow domains: 507 tasks across retail, auto insurance, travel, neobank, and consulting, with 477 graded on state alone and 30 adding response rubrics
- Failure taxonomy: The authors classify failures as tool usage 79.9%, wrong state updates 10.3%, incomplete user resolution 7.0%, and no state-changing action at all 2.9%, with most tool failures being retryable
- OpenEnv interface and local runs: The benchmark is served through OpenEnv and can be run locally on Linux or WSL with Python 3.11+, uv, and Docker; the same interface can feed training workflows, though the released adapter targets evaluation
Use Cases
- Engineering teams that need to verify an agent actually changes business state correctly before shipping it
- Model evaluators who care about consistency across repeated runs, where single-attempt pass rates and 20-of-20 rates produce two different leaderboards
- Teams studying recovery from tool errors, failed preconditions, and empty lookups
- Anyone costing agents per dependably completed task rather than per single success, which is the framing the authors use for their price comparison
Pros
- Grading rests on executable checks of terminal state, so 477 of the tasks do not depend on subjective scoring of the agent's reply
- The 20 independent reruns expose a consistency dimension: Kimi-K3 solved 93.89% of tasks at least once yet passed only 68 (13.41%) on all 20 attempts, a gap a single-attempt evaluation cannot see
- Harness under MIT and dataset under CDLA-Permissive-2.0, so both code and data are free to take and use
- Served through the OpenEnv interface, letting evaluation and training share the same environment definition
- The authors state plainly that tasks are synthetic reconstructions modeled on enterprise patterns with no real customers, and that cost figures come from an OpenRouter snapshot taken on September 20, 2026 at undiscounted list rates as a comparative index rather than cloud bills
Pricing
The harness is MIT-licensed and the benchmark dataset uses CDLA-Permissive-2.0, both available on Hugging Face, so running it yourself incurs no license fee. You supply your own model endpoints for the agent, simulated user, and judge roles; one endpoint can serve all three, which the authors call the simplest way to start. The environment requires Linux or WSL, Python 3.11+, uv, and Docker, plus a checkout of the dataset repository at the pinned release.
Summary
What makes ThinkingBox worth attention is that it moves the unit of grading from what the model said to what remains in the database. That shift produces hard numbers: across 121,680 valid trials on 12 models, 79,853 failed the executable checks, and 67.24% of those failures still terminated cleanly, having changed state, reported no final tool error, and looked entirely normal. On this measure none of the 18 models completed even half the tasks correctly on every attempt; Claude Opus 5 and 5.5 each managed 241 tasks (47.53%), with Opus 5.5 posting a higher single-attempt rate (67.16%) at the same dependable task count.
For people building agents, the failure taxonomy is more useful than the leaderboard: 79.9% of failures are classified as tool usage, and most of those are retryable failures rather than reasoning errors. The authors also caution that these are observable labels rather than causal explanations, and that they have not yet measured how much lift any of their recommended fixes would produce on the benchmark.
Version History
- ThinkingBox and ThinkingBox-Bench released (2026-10-03): Microsoft and Hugging Face released the ThinkingBox agent evaluation sandbox and the ThinkingBox-Bench benchmark, with 507 stateful enterprise workflows across retail, auto insurance, travel, neobank, and consulting, each run 20 times from a clean state and graded on terminal database state and side effects. Of 121,680 valid trials across 12 models, 79,853 failed, and 67.24% of those failures ended cleanly after a state-changing call with no reported tool error; among failures, 77.61% had wrong field values, 43.30% produced extra effects, and 25.36% omitted required effects. On the 18-model leaderboard Claude Opus 5.5 led single-attempt pass rate at 67.16%, while Kimi-K3 was the strongest open-weight model at 57.37% but passed only 68 tasks on all 20 attempts. The harness is MIT-licensed, the dataset uses CDLA-Permissive-2.0, and the benchmark is served through OpenEnv