The AI Post
Agents & CodingOpen ModelsEnterpriseFundraisingGenerative MediaGovernanceInferenceInfrastructureLegal & SafetySector Impact
← Front Page Research & Evals · OpenAI · Epoch AI · Anthropic

GPT‑6 Astra tops Epoch AI's IKEA assembly benchmark

OpenAI's latest model caught 80 percent of staged mistakes in Epoch AI's furniture-assembly test, up from 28 percent ten months ago, though it remains too slow for live guidance.

OpenAI's GPT-6 Astra scored 80 percent on a new Epoch AI test that checks whether a model can spot mistakes in photos of half-finished IKEA furniture, the research group said. That beats the 28 percent Claude Opus 4.5 managed when Epoch set the mark in November 2025.

The test uses 60 photos of three IKEA products: a shoe rack, a bed frame and a dresser, some showing a clean build and others hiding a planted mistake. Each model gets the printed manual, a zoom tool and a code interpreter, then has to say which step went wrong. Claude Fable 5.1 came second at 70 percent, and Claude Opus 5 scored 61 percent, according to Epoch AI.

Astra took a median of three minutes per photo, which Epoch AI called two to ten times faster than earlier leaders but still too slow to guide someone in real time. The group said the same skill could someday help with tasks like car repairs or appliance fixes. Grading used a separate model, GPT-5.6 Sol, to check each answer against the list of recorded mistakes.

Different models failed in different directions, Epoch AI said. Some flagged an error that was not there, and others missed a real one. The test covers just 60 photos across three products, so it is unclear whether the result would hold on a wider range of furniture tasks.

The group has not tested how a human scores on the same photos. It also found no Chinese open-weight model with vision good enough to compete — Kimi K3, the closest, still trails the frontier by about seven months.

Sources 2 sources

  1. Source Epoch AI
  2. Source The Decoder