Microsoft ThinkingBox grades AI agents by the database
Microsoft's ThinkingBox checks what an AI agent wrote to the database. 67.24% of 79,853 failed runs ended cleanly, with valid tool calls and wrong data.
3 min read

By the numbers
- tasks across five business domains
- 507
- runs of every task, to separate luck from reliability
- 20
- of failed runs ended cleanly with valid tool calls
- 67.24%
- of failures came from tool handling, not reasoning
- 79.9%
- Claude Opus 5.5
- 67.16%
- Claude Opus 5
- 66.5%
- GPT-5.4
- 65.36%
- GPT-6 Astra
- 58.31%
- Kimi-K3
- 57.37%
Microsoft released ThinkingBox, a test for AI agents that ignores what the agent says and checks what it actually changed. Microsoft announced it on October 3, 2026, in a Hugging Face blog post. Its central finding should worry anyone shipping agents. Of 79,853 failed attempts, 67.24% ended cleanly, with valid tool calls and no error. The data underneath was still wrong.
How ThinkingBox tests an agent
An AI agent is a model that takes actions through tools, such as updating an order or changing a booking. Most tests judge the agent's final reply or check that its tool calls were well formed. ThinkingBox looks at the back-end database instead, after the task ends.
RuntimeWire describes it as a benchmark where the database gets the final vote. Each of its 507 tasks simulates a real business workflow. Every task runs 20 times, so a lucky single success does not count as reliable.
The tasks cover five domains, according to Microsoft's post:
| Domain | Tasks |
|---|---|
| Retail | 98 |
| Auto insurance | 100 |
| Consulting (IT and HR support) | 101 |
| Travel | 104 |
| Neobank support | 104 |
The three scores
ThinkingBox reports three numbers per model, RuntimeWire explains. Pass@1 is the chance of success on a single attempt. Pass@20 means the task succeeded at least once in 20 runs. "Observed 20/20" counts the tasks the model passed every single time.
The gap between those numbers is the story. Claude Opus 5.5 led on pass@1 at 67.16%, but passed all 20 runs on only 47.53% of tasks, or 241 of them. Kimi-K3 scored 57.37% on pass@1 and passed all 20 runs on just 13.41%, or 68 tasks.
Microsoft's post says only three models kept more than 70% of their pass@1 score across the 20 trials.
Why agents fail quietly
Microsoft found that roughly four in five failures, 79.9%, came from tool handling rather than reasoning. The model understood the job but called a tool the wrong way, on the wrong record, or only partly.
That is how a run can look perfect and still be wrong. "An agent's final message is evidence of what it believes happened, not proof of what the software actually recorded," Remio writes. It warns that false confirmations on orders, subscriptions or permissions cause trouble later, when other systems act on them.
Cheap per attempt is not cheap per task
Microsoft also priced reliability. GPT-5.6 Sol costs $0.127 per successful attempt. Counted per task it completes reliably, the cost rises to $9.76. A model that is cheap per call can be expensive once you pay for the retries and the failures.
What this means for developers
Test your agent the way ThinkingBox does. After each test run, query the database and assert the exact fields that should have changed. Also assert that nothing else changed. A passing reply and clean tool logs prove very little on their own.
Run each test case many times, not once. A task that passes once can still fail often in production, as the Kimi-K3 numbers show. Track the share of tasks that pass every run, not only the average.
Spend your effort on the tool layer. If most failures come from tool handling, the fixes are often plain engineering. Write clearer tool descriptions, use stricter argument schemas, and return error messages the model can act on.
Add a check outside the model in production. Before telling a user that an order changed, read the order back from the system of record. Remio's point holds: the agent's own report is a claim, not proof.
Budget by reliable task, not by token. Microsoft's GPT-5.6 Sol figures show the per-call price can understate the real cost by a wide margin.
Sources
Related articles

Holo4 open-weight agent models score 61.7% on OSWorld 2.0
H Company's Holo4-27B scores 61.7% on OSWorld 2.0 at $1.22 a task, but only its Apache 2.0 sibling, 35B-A3B, can be self-hosted commercially.

NVIDIA Kumo Tabular tops TabArena with open weights
NVIDIA's Kumo Tabular predicts from a labeled table in one pass, with no training run. It ranks first on TabArena at 1,950 Elo, in sizes from 28M to 215M.

Liquid AI LFM2.5-VL-DSpark speeds up its 3B vision model
Liquid AI's 280M-parameter LFM2.5-VL-DSpark draft model decodes its 3B vision model up to 3.13x faster on an M5 Max, with identical output.
The daily brief
Three to five stories a day, and what each one means for the people who build software. Free, no spam.