OpenAI's GPT‑6 Astra Ultrafast generates tokens up to 8x faster, Nvidia says
Nvidia says a new Ultrafast mode of GPT-6 Astra, running on Blackwell GPUs, produces tokens up to eight times faster than Astra Standard and is live in the API.
Nvidia said on Thursday that GPT-6 Astra Ultrafast, a faster serving mode of OpenAI's GPT-6 Astra, is available now in the OpenAI API and to eligible ChatGPT Work and Codex users. The company says it offers up to eight times faster token generation than Astra Standard, and runs on Nvidia Blackwell GPUs.
The speed figure comes from Nvidia's own blog post, which describes the mode as aimed at coding, tool use and interactive applications. It gives no benchmark method, no baseline hardware and no latency numbers, and the phrase up to means typical gains may be lower. OpenAI's own announcement was not part of the material we reviewed.
The post does not state a price, and sends readers to an Ultrafast guide for it. Faster serving modes usually cost more per token, but the size of any premium here is not confirmed. Developers should check the guide before moving production traffic onto the new mode.
Nvidia quotes two OpenAI executives. Philippe Tillet, its inference lead, said Nvidia's tooling and documentation helped OpenAI make its models exceptionally good at programming Blackwell. Uday Ruddarraju, its chief technology officer of compute, said OpenAI used internal models to optimise inference on Nvidia GPUs, and that this allows continuing gains after launch.