Cognition is first to run Nvidia Vera Rubin on CoreWeave, claims 4.8x
Cognition says Nvidia's Vera Rubin NVL72 racks deliver up to 4.8 times the token throughput of GB200 on its SWE-2 inference, and that Devin workloads moved within days of delivery.
Cognition, the company behind the Devin coding agent, said it is the first customer to run production work on Nvidia's Vera Rubin NVL72 systems, hosted by CoreWeave. In a post on X, Cognition said that on its SWE-2 inference workload the new chips deliver about 4.8 times the token throughput of GB200 at the same decode speed.
https://x.com/cognition/status/2105408461701824732
CoreWeave's announcement, made at its Fully Connected conference, gives the same 4.8 times figure and adds a second: 3.8 times the output-token throughput for reinforcement learning. The system entered limited availability on CoreWeave Cloud in September. CoreWeave said Cognition put production workloads on the racks within days of delivery, without custom infrastructure engineering.
Silas Alberti, Cognition's senior vice president of research, said engineers are seeing up to a 4.8 times increase in total token throughput, which he said could let Devin finish work faster at lower cost per token. Both figures come from the two companies, and an outside party may yet test them.
CoreWeave lists each NVL72 rack as 72 Rubin GPUs linked by NVLink 6, plus 36 Vera CPUs, with 1,400 terabytes a second of HBM4 memory bandwidth. It is liquid cooled and is designed to act as a single logical GPU. Cognition said it runs training, reinforcement learning and production inference on the same bare-metal infrastructure.