Artificial Analysis says top AI models refuse third of cyber tasks
Artificial Analysis's new Cyber Index, built with Collinear, IBM, Nvidia and Vercel, found Grok 4.7 and MiMo-V2.6-Pro on top while Claude and GPT-6 models refused a third of tasks.
Artificial Analysis launched a Cyber Index and Cyber Index Alliance with Collinear AI, IBM, Nvidia and Vercel, a new benchmark measuring how well AI agents find and patch vulnerabilities in code.
https://x.com/ArtificialAnlys/status/2104548886442647864
The index combines three benchmarks. CWE-Bench-AA, from Collinear, tests patching across 120 tasks spanning the OWASP Top 10. DeepsecBench-AA, from Vercel, scores vulnerability discovery against a set of expert-verified findings. CyberGym-E2E-AA, from Berkeley RDI, asks agents to find a memory-safety bug, write a proof-of-concept exploit, then patch it.
Grok 4.7 and MiMo-V2.6-Pro topped the leaderboard with a score of 56, Artificial Analysis said. GPT-6 Luna followed at 53, then GLM-5.3-Flash at 50 and Muse Spark 1.3 at 44.
Several frontier models scored lower because they refused tasks on safety grounds, Artificial Analysis said. Claude Opus 5.5, Claude Fable 5.1, Gemini 3.8 Flash and two GPT-6 variants declined 32% to 38% of the index's tasks. That left them trailing the leaders by 19 to 31 points.
Most of that gap came from CyberGym-E2E-AA, the end-to-end benchmark, Artificial Analysis said. GPT-6 Sol and GPT-6 Astra refused every task there. Claude Opus 5.5 refused 98% of tasks and Claude Fable 5.1 refused 99%. The company did not say why refusal rates varied so widely by benchmark.