GPT‑6 Astra attempts 97 of 100 unsafe robot instructions, Robocurve says
The benchmark gave three policies five instructions a safe robot should refuse, and only Anthropic's Claude Fable 5.1 turned down any of them every time.
The benchmark gave three policies five instructions a safe robot should refuse, and only Anthropic's Claude Fable 5.1 turned down any of them every time.
The company says it will track how much of its own research is done by AI, how well agents are overseen, and how compute is allocated.
The company reports a fourfold speed-up and is funding a protein design competition to test more than five thousand designs experimentally.