2026-10-03 · ← News
ThinkingBox grades agents by the database, not by confident answers
Microsoft and Hugging Face have made ThinkingBox available, a benchmark of 507 stateful business tasks that runs each task 20 times and inspects the actual backend outcome. It shows why a successful tool call or polished answer does not mean the job was completed.
The image could not be loaded.
Microsoft and Hugging Face have made ThinkingBox available, grading an agent by the records and side effects it leaves in the backend. The benchmark shifts attention from conversational plausibility to an outcome that can be checked.
507 tasks end with an inspection of actual state
ThinkingBox contains 507 executable tasks across retail, auto insurance, travel, neobanking and consulting support. Every attempt starts from the same clean state. A simulated user supplies missing details, and tests compare the final database with the expected outcome.
The authors repeat every task 20 times. In an analysis of 121,680 valid trials across 12 models, 79,853 runs failed. Yet 67.24% of those failures ended without a reported tool error and included a state changing call. The agent often looked successful exactly when the database said otherwise.
Production teams need to test outcomes and repeatability
For teams deploying agents, capability and reliability are different questions. A model may solve a task once without being safe across hundreds of cases. ThinkingBox therefore reports pass@1 alongside the tasks that pass all 20 attempts.
The practical consequence is specific. Evals must inspect the correct record, forbidden side effects and the state left after a tool failure. A transcript helps diagnose the path, but it is not proof that a transaction finished correctly.
Controlled sandboxes still simplify production
The benchmark uses synthetic data, bounded MCP tools and a known initial state. That makes changes attributable to one run, but it excludes much of the mess in live systems, including concurrent writes and unexpected integrations. Its score is a measurement, not a warranty for a deployment.
Its cost estimates also use recorded tokens and undiscounted OpenRouter list prices. The authors call them a comparative index rather than a production bill.
Errors caught before customers see them will decide its value
The next useful signal will come from teams that bring final state checks into their own evals. The question is whether repeated runs expose bad refunds, prematurely closed tickets and unintended writes before deployment.
ThinkingBox is available through two repositories, separating the runtime from the data, synthetic records and MCP servers. That provides a usable foundation, but the real work begins when a team writes assertions for its own workflows.
Lilith's verdict
An agent declaring success before checking the database is a courier holding a signed slip while the parcel remains in the van. ThinkingBox finally grades the parcel, not the signature.
I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.
Original source ↗ ↗