Skip to content

Microsoft ThinkingBox grades AI agents by the database

Microsoft's ThinkingBox checks what an AI agent wrote to the database. 67.24% of 79,853 failed runs ended cleanly, with valid tool calls and wrong data.

By Tech AI Wire Team

3 min read

XLinkedIn
A screenshot of Microsoft's ThinkingBox-Bench dataset page on Hugging Face, with the dataset viewer listing retail and e-commerce tasks.

By the numbers

tasks across five business domains
507
runs of every task, to separate luck from reliability
20
of failed runs ended cleanly with valid tool calls
67.24%
of failures came from tool handling, not reasoning
79.9%
ThinkingBox pass@1 (single-attempt success)
Claude Opus 5.5
67.16%
Claude Opus 5
66.5%
GPT-5.4
65.36%
GPT-6 Astra
58.31%
Kimi-K3
57.37%

Microsoft released ThinkingBox, a test for AI agents that ignores what the agent says and checks what it actually changed. Microsoft announced it on October 3, 2026, in a Hugging Face blog post. Its central finding should worry anyone shipping agents. Of 79,853 failed attempts, 67.24% ended cleanly, with valid tool calls and no error. The data underneath was still wrong.

How ThinkingBox tests an agent

An AI agent is a model that takes actions through tools, such as updating an order or changing a booking. Most tests judge the agent's final reply or check that its tool calls were well formed. ThinkingBox looks at the back-end database instead, after the task ends.

RuntimeWire describes it as a benchmark where the database gets the final vote. Each of its 507 tasks simulates a real business workflow. Every task runs 20 times, so a lucky single success does not count as reliable.

The tasks cover five domains, according to Microsoft's post:

DomainTasks
Retail98
Auto insurance100
Consulting (IT and HR support)101
Travel104
Neobank support104

The three scores

ThinkingBox reports three numbers per model, RuntimeWire explains. Pass@1 is the chance of success on a single attempt. Pass@20 means the task succeeded at least once in 20 runs. "Observed 20/20" counts the tasks the model passed every single time.

The gap between those numbers is the story. Claude Opus 5.5 led on pass@1 at 67.16%, but passed all 20 runs on only 47.53% of tasks, or 241 of them. Kimi-K3 scored 57.37% on pass@1 and passed all 20 runs on just 13.41%, or 68 tasks.

Microsoft's post says only three models kept more than 70% of their pass@1 score across the 20 trials.

Why agents fail quietly

Microsoft found that roughly four in five failures, 79.9%, came from tool handling rather than reasoning. The model understood the job but called a tool the wrong way, on the wrong record, or only partly.

That is how a run can look perfect and still be wrong. "An agent's final message is evidence of what it believes happened, not proof of what the software actually recorded," Remio writes. It warns that false confirmations on orders, subscriptions or permissions cause trouble later, when other systems act on them.

Cheap per attempt is not cheap per task

Microsoft also priced reliability. GPT-5.6 Sol costs $0.127 per successful attempt. Counted per task it completes reliably, the cost rises to $9.76. A model that is cheap per call can be expensive once you pay for the retries and the failures.

What this means for developers

Test your agent the way ThinkingBox does. After each test run, query the database and assert the exact fields that should have changed. Also assert that nothing else changed. A passing reply and clean tool logs prove very little on their own.

Run each test case many times, not once. A task that passes once can still fail often in production, as the Kimi-K3 numbers show. Track the share of tasks that pass every run, not only the average.

Spend your effort on the tool layer. If most failures come from tool handling, the fixes are often plain engineering. Write clearer tool descriptions, use stricter argument schemas, and return error messages the model can act on.

Add a check outside the model in production. Before telling a user that an order changed, read the order back from the system of record. Remio's point holds: the agent's own report is a claim, not proof.

Budget by reliable task, not by token. Microsoft's GPT-5.6 Sol figures show the per-call price can understate the real cost by a wide margin.

Sources

  1. The Agent Said It Was Done. The Database Disagreed. - Hugging Face Blog
  2. Microsoft releases an agent benchmark where the database gets the final vote - RuntimeWire
  3. Microsoft ThinkingBox Benchmark Exposes the Gap Between Agent Claims and Database Reality - Remio

Related articles

The daily brief

Three to five stories a day, and what each one means for the people who build software. Free, no spam.