ChengRang

Microsoft ThinkingBox

AI Search & Research Open Source

An agent evaluation sandbox from Microsoft and Hugging Face, running 507 stateful enterprise workflows 20 times each and grading the terminal database state and side effects instead of what the agent said

Agent EvaluationBenchmarkState GradingMCPOpen Source
Visit Microsoft ThinkingBox

Disclaimer: Review content represents our editorial team's views and experience, not commercial recommendation or investment advice. Product info and pricing may change; refer to official sources.

Overview

ThinkingBox is an agent evaluation sandbox released jointly by Microsoft and Hugging Face on October 3, 2026, together with a benchmark named ThinkingBox-Bench. Its grading principle is blunt: ignore what the agent said at the end, ignore how many tool calls it made, and look at what is left in the backend database when the run is over.

The 507 tasks span five enterprise workflow domains: retail, auto insurance, travel, neobank, and consulting. Each task defines a starting backend state, a user goal, the available MCP tools, and a domain policy, then runs 20 independent attempts from an identical clean state. Every attempt gets an isolated MCP session, and no two attempts share a database row or cached tool state, which is what makes a 20-trial comparison meaningful. Afterwards a side-effect extractor derives what actually changed, and deterministic judges compare it against the required end state: any trajectory producing the right outcome passes, while wrong, missing, or extra effects fail. Of the 507 tasks, 477 are graded on state alone and 30 add a binary rubric for requirements with no clean database value, such as whether the agent disclosed that a result is not guaranteed.

The motivating example from the authors makes the point well: a support agent handles a late delivery, makes nine tool calls that all look correct, reads the refund policy correctly, and closes the ticket as resolved. Two things are still wrong: the courier exception remains open and the customer's actual question was never answered. A grader checking only tool calls or the final sentence would score it as a success. The database disagrees.

Both the sandbox and the dataset are on Hugging Face. The harness is MIT-licensed, the dataset uses CDLA-Permissive-2.0, the benchmark is served through the OpenEnv interface, and each finished episode returns a binary pass or fail reward.

Key Features

Use Cases

Pros

Pricing

The harness is MIT-licensed and the benchmark dataset uses CDLA-Permissive-2.0, both available on Hugging Face, so running it yourself incurs no license fee. You supply your own model endpoints for the agent, simulated user, and judge roles; one endpoint can serve all three, which the authors call the simplest way to start. The environment requires Linux or WSL, Python 3.11+, uv, and Docker, plus a checkout of the dataset repository at the pinned release.

Summary

What makes ThinkingBox worth attention is that it moves the unit of grading from what the model said to what remains in the database. That shift produces hard numbers: across 121,680 valid trials on 12 models, 79,853 failed the executable checks, and 67.24% of those failures still terminated cleanly, having changed state, reported no final tool error, and looked entirely normal. On this measure none of the 18 models completed even half the tasks correctly on every attempt; Claude Opus 5 and 5.5 each managed 241 tasks (47.53%), with Opus 5.5 posting a higher single-attempt rate (67.16%) at the same dependable task count.

For people building agents, the failure taxonomy is more useful than the leaderboard: 79.9% of failures are classified as tool usage, and most of those are retryable failures rather than reasoning errors. The authors also caution that these are observable labels rather than causal explanations, and that they have not yet measured how much lift any of their recommended fixes would produce on the benchmark.

Version History

Category
AI Search & Research
Pricing
Open Source
Tags
Agent Evaluation · Benchmark · State Grading
Website

Related Tools