The AI Post
Agents & CodingOpen ModelsEnterpriseFundraisingGenerative MediaGovernanceInferenceInfrastructureLegal & SafetySector Impact
← Front Page AI · cua-speedrun · GPT-6 Astra · Claude Opus 5 · Gemini 3.8 Flash

Agent speed benchmark puts GPT‑6 Astra first, open models nowhere

A new benchmark times AI agents on desktop tasks. Its authors report GPT-6 Astra scoring 90.8% on OSWorld in 90 seconds at $0.71 a task, while Gemini 3.8 Flash matched the top score at $0.11.

Researchers including Jing Yu Koh, Ruslan Salakhutdinov and Daniel Fried published cua-speedrun on arXiv on 30 September, a standardized way to measure how fast computer-use agents finish graphical-interface tasks. The authors argue that slow, costly runs limit deployment and that existing benchmarks are hard to reproduce. They run everything on identical Modal cloud infrastructure and time each task from instruction to finish.

The study covers 56 agent configurations on OSWorld and 21 on OSWorld2, across four benchmarks. On a 50-task OSWorld subset, the authors report GPT-6 Astra at low effort reaching 90.8% accuracy in 90 seconds at $0.71 per task, and Claude Opus 5 at 87.6% in 86 seconds. Gemini 3.8 Flash matched the top score in similar time at $0.11 per task.

The authors say no single model family is best on all three measures at once, though GPT-6 Astra configurations alone form the performance-time frontier on OSWorld2. None of the open-weight models they tested, including Kimi K3, MiniMax M3 and GLM-5V Turbo, sat on any performance-time or performance-cost frontier, according to the paper.

One finding runs against intuition. On Gemini Flash, moving from low to medium reasoning effort raised the score and cut mean task time from 492 to 268 seconds, which the authors attribute to fewer repeated failed actions. Higher effort on OSWorld added time and cost without gains. The authors caution that results reflect API speeds and prices at test time, which providers can change.

Sources 2 sources

  1. Source arXiv
  2. Source cua-speedrun