A custom ASIC to rule AI efficiency domestically?
When you purchase through links on our site, we may earn an affiliate commission. Here’s how it works.
OpenAI took the stage at Hot Chips on August 25 with the first published performance figures for Jalapeño, the inference accelerator it co-developed with Broadcom, and the numbers are designed to be read the way it flatters the former: efficiency.
The company reported 1.5 to 1.9 times more throughput per kilowatt and 1.7 to 3.6 times lower end-to-end latency across three open models than the Nvidia rack systems it tested against.
OpenAI's Jalapeño is currently rated at 700W, versus comparable Nvidia silicon that was rated for 1200W and 1400W, a key differentiator in a benchmark that already has limited information available to researchers looking to pick a clear winner.
OpenAI's Jalapeño is an ASIC, or an Application-specific integrated circuit, which essentially means that it focuses primarily on and is great for very specific AI workloads; in this case, OpenAI's inference needs, which allow it to run multiple models with significant efficiency gains in tow.
Richard Ho, who runs OpenAI's hardware program, told reporters on a press call that the results show "a very, very significant performance advance over state-of-the-art."
OpenAI ran SemiAnalysis's public InferenceX suite on GPT-OSS 120B, DeepSeek R1 670B, and Moonshot AI's trillion-parameter Kimi K2.5, at a nominal 8,000-token input and 1,000-token output. It scored a 1.9x efficiency win vs. GB200 on its own GPT-OSS model, and 1.7x and 1.5x on DeepSeek and Moonshot's offerings against a GB300 system. Interestingly, the figures for GPT-OSS 120B running on a GB300 compared to OpenAI's Jalapeño are not published.
It is important to point out here that there is a certain degree of cherry-picking to flatter OpenAI's own results: by choosing to compare ratios on a per-kilowatt basis versus a per-chip basis, while efficiency remains consistent, it does paint Jalapeño as potentially a much more potent competitor to Nvidia's last-generation offerings than it actually is.
Sign up to the TechRadar Pro newsletter to get all the top news, opinion, features and guidance your business needs to succeed!
Nvidia's GB300 might not have an efficiency win versus Jalapeño, but it remains a much more powerful chip in all tests, at a time when Nvidia is already rolling out Vera Rubin. Curiously, OpenAI has not published any information about how its chip stacks up against an HBM4-touting Vera Rubin, which Nvidia has presented as an efficiency juggernaut.
It is important to point out that Jalapeño's results also use single-token prediction throughout, with no speculative decoding and no prefill-decode disaggregation. Nvidia's production deployments commonly use multi-token prediction, and when OpenAI pits Jalapeño against a GB300 configured that way, the peak efficiency lead drops to roughly 1.5x.
OpenAI says it built the chip to keep model state local, writing that it "designed Jalapeño to minimize data movement and communication delays." Cores and HBM are divided into slices, each core slice holding a low-latency view of its own memory, with synchronization pushed onto a dedicated collective network. Each package pairs a compute die with six HBM4 stacks for 216 GiB at 15.4 TB/s, which is less capacity than the GB300's 288GB of HBM3E but considerably more bandwidth, and roughly 50 percent more memory per watt of rated power.
Every number published currently comes from A0 silicon, also known as the initial, first-run version of the chip. A more efficient B0 stepping is already in the fab with roughly 25 percent better performance per watt, and production is scheduled to ramp gradually across 2027.
Although it may not hold a candle to Nvidia's Vera Rubin, it might not have to; Alexander Harrowell of Omdia told CNBC that "this is the biggest competitive threat to NVIDIA," noting that roughly half of AI infrastructure capital spending comes from firms that already run custom chip programs or could. Yole Group's Adrien Sanchez framed it as pressure on Nvidia's inference margins specifically, the fastest-growing part of the business.


