Microsoft's ThinkingBox checks what an AI agent actually changed in a database, not what it said. Across 121,680 trials, the authors say 67% of failures were clean tool calls with wrong values.
Microsoft's Copilot Studio team and the data company Toloka published ThinkingBox on 3 October, a benchmark that scores AI agents on the state they leave in a database rather than on their replies. The authors say that across 121,680 trials, 67.24 per cent of failures still called tools cleanly but wrote wrong field values. No outside group has reproduced the numbers.
Continue reading