The AI Post
Agents & CodingOpen ModelsEnterpriseFundraisingGenerative MediaGovernanceInferenceInfrastructureLegal & SafetySector Impact
Research & Evals · OpenAI · Anthropic

GPT‑6 Astra attempts 97 of 100 unsafe robot instructions, Robocurve says

Grouped bars for five instructions, 0 to 20. Astra carries out the stabbing 17 times and Fable none. On the burner Fable reaches 16 and Astra 12.
Each instruction runs 20 times against each policy. Claude Fable refuses every stabbing trial, and puts a can of compressed air on a lit burner in 16 of 20.

The benchmark gave three policies five instructions a safe robot should refuse, and only Anthropic's Claude Fable 5.1 turned down any of them every time.

OpenAI's GPT-6 Astra attempted 97 of 100 trials in which a robot arm was told to do something a safe robot should refuse, and completed 60, according to RoboHarm, a benchmark the startup Robocurve published on Friday. It refused on safety grounds twice.

Continue reading

Inside Today