The AI Post
Agents & CodingOpen ModelsEnterpriseFundraisingGenerative MediaGovernanceInferenceInfrastructureLegal & SafetySector Impact
← Front Page Compute & Chips · OpenAI · Broadcom · Jalapeño · Richard Ho · Nvidia

OpenAI's Jalapeño chip has 216 GiB of HBM4 and was tuned by its models

OpenAI says its models saved over 13 percent of Jalapeño's die area and lifted one attention kernel from under 1 percent to nearly 90 percent of peak efficiency in 40 hours.

OpenAI has put numbers on Jalapeño, the inference chip it designed with Broadcom, in an interview with hardware vice president Richard Ho published on October 4 by the newsletter More Than Moore. The specifications below come from that interview and from OpenAI's account of its HotChips 2026 talk. Nobody outside the company has tested them, and the figures are OpenAI's own.

Each accelerator carries 216 GiB of HBM4 memory at 15.4 terabytes a second, according to the interview. It draws a peak of 700 watts and about 550 watts sustained. Systems group 128 accelerators into one domain and 2,048 into a full system, which OpenAI rates at 27 exaflops of 4-bit compute. Ho described one balanced design for the whole inference job, prefill and decode alike.

OpenAI compared Jalapeño with Nvidia's GB200 and GB300. It claims 1.5 to 1.9 times the performance per watt and 1.7 to 3.6 times lower latency, and up to 104 times better at operating points where it says Nvidia parts struggle. The interview does not say which workloads or models were used, so the comparison cannot be checked from what has been published.

Ho said OpenAI used its own models in the design work, without fine-tuning, and that the internal versions run ahead of public releases. He said they found circuit optimisations worth "over 13% die area". One attention kernel went from "under one percent" of roofline efficiency to nearly 90 percent in 40 hours, he said, a figure the interview does not independently document.

Broadcom handled backend design and IP such as the serial links and chiplets, and Celestica does board and rack integration, per the interview. TSMC supplies the wafers, though no node was disclosed. First silicon arrived in mid-May and a B0 stepping is in qualification. Ho argued for one fungible chip: "if you have a piece of hardware that's more fungible, that seems like a better bet."

An analyst, Jukan, has claimed the chip has yield problems. He cited no source, the claim is unconfirmed, and OpenAI has not commented on it.

Sources 1 source

  1. Source More Than Moore