The AI Post
Agents & CodingOpen ModelsEnterpriseFundraisingGenerative MediaGovernanceInferenceInfrastructureLegal & SafetySector Impact
Research & Evals · Microsoft · Toloka · ThinkingBox · Claude Opus 5.5 · Kimi-K3 · GPT-6 Astra · Hugging Face

Microsoft benchmark finds agents often claim success but leave databases wrong

Microsoft's ThinkingBox checks what an AI agent actually changed in a database, not what it said. Across 121,680 trials, the authors say 67% of failures were clean tool calls with wrong values.

Microsoft's Copilot Studio team and the data company Toloka published ThinkingBox on 3 October, a benchmark that scores AI agents on the state they leave in a database rather than on their replies. The authors say that across 121,680 trials, 67.24 per cent of failures still called tools cleanly but wrote wrong field values. No outside group has reproduced the numbers.

Continue reading

Snippets
  • The X account thsottiaux, which posts about OpenAI's Codex, replied "6.1 coming soon" to a user on Sunday; the post has more than 1,800 likes. It did not say what 6.1 refers to. thsottiaux on X
  • Commentator kimmonismus wrote on X that Gemini 4 is "very close to official release" and that GLM‑5.3 is still cited in Anthropic blog posts as among the best models for cyber capabilities. kimmonismus on X
  • A Meta paper called RankEvolve runs Claude Code and Codex as separate nodes that review each other's changes; Meta authors say the pair reach 62.5% execution accuracy against 45.8% for the best single product. dair_ai on X
  • A CMU paper on harness learning trains a 4‑billion‑parameter proposer with reinforcement learning to edit an agent's harness code, and omarsar0 says it beats its 35B teacher at single‑step revision on Reasoning Gym. omarsar0 on X
  • Andrew Curran says the White House's AI task force has been renamed the SI Force and will be led by Jay Clayton, citing the WSJ, with no official announcement yet. Andrew Curran on X
  • 9to5Google reports that from October 9 the Gemini app's free tier loses Flash and Pro and keeps only Flash‑Lite, while $4.99 AI Plus loses Pro and the $19.99 AI Pro tier gains Deep Think. 9to5Google
  • Kevin Liao argues in a blog post that agent memory plugins are all retrieval over stored snippets, and that agents need written project documentation instead; the post has 61 points on Hacker News. Kevin Liao
  • Tom's Hardware reports Google has frozen product‑vulnerability submissions to its open‑source bug bounty program until 2027 because maintainers are drowning in invalid, AI‑generated reports. Tom's Hardware
  • Tom's Hardware reports a Futurum chief says AI agents use five times as many tokens as humans, heading for ten times, mostly because cached prompts are re‑read repeatedly rather than new work. Tom's Hardware
All snippets