The AI Post
Agents & CodingOpen ModelsEnterpriseFundraisingGenerative MediaGovernanceInferenceInfrastructureLegal & SafetySector Impact
← Front Page Research & Evals · Microsoft · Toloka · ThinkingBox · Claude Opus 5.5 · Kimi-K3 · GPT-6 Astra · Hugging Face

Microsoft benchmark finds agents often claim success but leave databases wrong

Microsoft's ThinkingBox checks what an AI agent actually changed in a database, not what it said. Across 121,680 trials, the authors say 67% of failures were clean tool calls with wrong values.

Microsoft's Copilot Studio team and the data company Toloka published ThinkingBox on 3 October, a benchmark that scores AI agents on the state they leave in a database rather than on their replies. The authors say that across 121,680 trials, 67.24 per cent of failures still called tools cleanly but wrote wrong field values. No outside group has reproduced the numbers.

The write-up, posted on Hugging Face, runs 18 models against 507 business-workflow tasks inside isolated tool sessions. Each task is attempted 20 times, because the authors argue that one success is not reliability. Their stated thesis is that a tool call is not an outcome. Collaborators include researchers at the University of Pittsburgh and Northwestern University.

On a single attempt, Anthropic's Claude Opus 5.5 leads at 67.16 per cent, the authors report. Moonshot's Kimi-K3 solves 93.89 per cent of tasks at least once but only 13.41 per cent every time. Opus 5.5 and Opus 5 each pass 241 tasks on all 20 tries, the most consistent results in the table.

The authors also price dependability. They say GPT-5.4 has the lowest cost per dependable task at $6.80, while GPT-6 Astra reaches more consistent results, 231 tasks, at $7.45 a task. Roughly four in five failures come from tool handling rather than reasoning, according to the write-up, which suggests the weak point is plumbing, not intelligence.

The framework is MIT-licensed and the benchmark data uses the CDLA-Permissive-2.0 licence. Both are on Hugging Face as microsoft/ThinkingBox-Bench and as an OpenEnv environment under BSD-3-Clause. Because the results come from the benchmark's own authors, and model makers have not commented, the rankings should be read as one team's measurement until others rerun them.

Sources 1 source

  1. Source Hugging Face (Microsoft)