Independent evaluators split on what Grok 4.7 improved
Artificial Analysis put the model in the top four on its intelligence index, while Vals AI measured a decline driven by finance and research tasks.
Two independent evaluation outfits published results for SpaceXAI's Grok 4.7 on Monday and disagreed about whether the model is an improvement on Grok 4.6. Both ran the model at its highest reasoning setting, and both published their scores in full.
Artificial Analysis scored Grok 4.7 at 46 on its Intelligence Index, which it said brings SpaceXAI into the top four AI labs. Grok Build running Grok 4.7 scored 56 on its Coding Agent Index, up from 47 with Grok 4.6, and the firm put the hallucination rate at 29 percent against 34 percent.
Vals AI read its own numbers the other way. It said most of the decline in its index comes from finance and research-heavy tasks, with its EMB test falling 7.8 points and Finance Agent v2 falling 4.5 points, together about two-thirds of the drop.
The two ran different tests and report different things, so the results may not be comparable. Vals AI said it used SpaceXAI's default provider settings at xhigh reasoning. Artificial Analysis also tested at xhigh, and put output speed at about 188 tokens a second on long prompts.
Both measured the model consuming more. Artificial Analysis said Grok 4.7 used 81,000 output tokens per Intelligence Index task, more than double the 36,000 used by Grok 4.6. Vals AI put the cost per test at $4.78 against $4.34, on roughly 15 percent more input tokens.
Elon Musk said Grok 4.7 places SpaceXAI third for agentic coding, behind Anthropic and OpenAI, and called the model "a strong combination of intelligence, speed & low cost". That ranking is the company's own reading of the published indices.