The AI Post
Agents & CodingOpen ModelsEnterpriseFundraisingGenerative MediaGovernanceInferenceInfrastructureLegal & SafetySector Impact
← Front Page Research & Evals · Artificial Analysis · OpenAI · Anthropic · Terminal-Bench

Artificial Analysis says only 2 models clear 50% on science benchmark

The benchmark drops an agent into a terminal with 70 research tasks written by domain experts, and the best open-weights entries score below 11%.

Artificial Analysis launched a leaderboard for Terminal-Bench-Science 0.1 on Thursday night, and said only two models pass more than half of its 70 tasks. It has run 26 of the 33 models on the page so far.

GPT-6 Astra at its maximum reasoning setting scores 63.3%, according to the leaderboard, with Claude Opus 5.5 on 61.9% at its xhigh setting and 59.0% at max. Artificial Analysis said it ran the evaluations itself, using the mini-swe-agent harness.

The open-weights entries sit far behind. Artificial Analysis put GLM-5.3 at 10% and DeepSeek V4.1 Flash at 9%, more than 50 points below the leaders. It said recent releases have moved the top scores a long way, and that GPT-6 Astra still has room above it.

Domain experts wrote and reviewed the tasks, which split 19 in life sciences, 17 in physical sciences, 17 in mathematical sciences, 9 in engineering and 8 in earth sciences. Each task runs in a terminal, and its own tests grade it pass@1 over three attempts.

Artificial Analysis said the domain scores may be noisy, because some rest on as few as 8 tasks. It put Claude Opus 5.5 at xhigh on 71% of the mathematical sciences tasks and 46% of the life sciences tasks. Steven Dillmann and researchers at Stanford built the benchmark with the Terminal-Bench team, and announced it in August.

Sources 2 sources

  1. Source Artificial Analysis
  2. Source Artificial Analysis on X