The AI Post
Agents & CodingOpen ModelsEnterpriseFundraisingGenerative MediaGovernanceInferenceInfrastructureLegal & SafetySector Impact
← Front Page Research & Evals · Agent Arena · xAI · SpaceX · Anthropic · OpenAI

Agent Arena puts Grok 4.7 16th, at $1.13 per task

The leaderboard gives the model a net improvement score of 3.96% over 6,801 sessions, against 13.44% for the Claude Fable 5.1 entry on top.

Agent Arena placed Grok 4.7 16th on its leaderboard early on Friday, with a net improvement score of 3.96% and a median cost of $1.13 a task. Agent Arena credited the model to SpaceX, and said it ran at the model's xHigh reasoning setting.

Agent Arena said Grok 4.7 improves on earlier versions at a comparable lift in price. It put Grok 4.6 at its high setting on 1.22 net improvement and $0.74 a task, and Grok 4.5 on 1.50 and $0.43. Both scored below the new entry.

By signal, Agent Arena ranked Grok 4.7 sixth on confirmed success at 10.41%, 12th on bash recovery at 5.98% and 16th on praise against complaint at 4.93%, according to its post. It put the model 28th on steerability, at -1.89%, and said tool hallucination showed no issue.

The entries above it cost more. The leaderboard puts Claude Fable 5.1 at its max setting first, on 13.44% and $4.15 a task. It puts GPT-6 Astra at max second, on 11.08% and $2.85, and Claude Opus 5 at high third, on 9.82% and $2.15.

The figures are Agent Arena's own measurements, taken over 6,801 sessions, and it gave the Grok 4.7 score a margin of 2.07 percentage points. No one else has published the same comparison of the three Grok releases, and the post did not say how net improvement is computed.

Sources 2 sources

  1. Source Agent Arena leaderboard
  2. Source Agent Arena on X