Engineer says Opus 5.5 wrote a GPU kernel at 35.5 times PyTorch
Elliot Arledge published a run-by-run comparison on his own hardware and put GPT-6 Astra second at 24.8 times the baseline.
A performance engineer said Claude Opus 5.5 wrote a GPU kernel that ran 35.5 times faster than an optimised PyTorch baseline on the KernelBench-Mega benchmark, against 24.8 times for GPT-6 Astra. Elliot Arledge posted the result on Wednesday, describing the output as a Kimi-Linear decode megakernel.
Arledge said he gave each model one session on his own RTX PRO 6000 machine, with no budget cap. The figures are isolated regrades, he said, made after an audit of the run traces and submissions for reward hacking. That check matters: a kernel benchmark can be gamed by a model that weakens the baseline it is timed against.
The runs differed more in shape than in their final numbers. Arledge said Opus 5.5 had a working megakernel at 21 times the baseline after 31 minutes. Astra reached 14 times after an hour, then spent four hours moving from 24 to 25. Fable 5.1 reached 23 times in two hours before the session died on an API error.
One session per model is a small sample and the traces are not published, so the ordering may not hold on a rerun. Arledge said he runs 20 such jobs in parallel. That is why he values how the model reports status: he said it takes him seconds to see what is blocking a run. He also said he suspects reports of weakened kernel performance apply to TPU or Trainium rather than CUDA.
A week of testing at Every produced a more mixed account. The publication said Opus 5.5 will not stop on its own and will eat a weekly limit. Given ten minutes to produce a run of show, it built a data generator and handouts, then ran out of time before the schedule. It also rejected 90% of its own work for reasons its designer could not trace.
Every also said the model builds on feedback instead of arguing with it. Its prose is the most readable the publication has measured, it said, though it still buries the lede. On cost, the account SahilExec ran the same prompt through the same harness. It logged $75 and 2 hours 10 minutes for Opus 5.5, against $61 and 1 hour 25 minutes for Astra.