Skip to content
Live newsroom 148 readers online
Tuesday, August 25, 2026 Live Sync: Just now
Demystifying Finance, Technology, and Global Markets for the Next Generation.
BreakingNvidia releases DLSS 4.5 Ray Reconstruction, available now in 30 games
Share Suggestions AVOID NVDA Stage 4 (Conv: 1/5 | Size: 10%)

OpenAI's upcoming Jalapeño chip looks like it'll be an inference beast

128 chips, 1.7 exaFLOPS, and 27 TB of HBM give Altman and crew a leg up over Blackwell, and maybe even Rubin OpenAI offered its closest look yet at its spicy new Jalapeño AI accelerator at the annual Hot Chips semiconductor development conference at Stanford on Tuesday. The chips, first teased earlier this year, were […]

By deepak · August 25, 2026 · 3 min read

128 chips, 1.7 exaFLOPS, and 27 TB of HBM give Altman and crew a leg up over Blackwell, and maybe even Rubin

OpenAI offered its closest look yet at its spicy new Jalapeño AI accelerator at the annual Hot Chips semiconductor development conference at Stanford on Tuesday.

The chips, first teased earlier this year, were developed in collaboration with Broadcom, and are the first in a series of custom silicon from OpenAI, designed (in part) by AI, for AI.

Compared to contemporary GPU systems from Nvidia, OpenAI says the parts will deliver both higher throughput and lower latency when they start trickling out later this year and reach volume production in 2027.

To be clear, Jalapeño won’t replace OpenAI’s long-time hardware partners, which also happen to be some of its most important investors.

OpenAI still needs compute for training, and the highly programmable nature of GPUs means that OpenAI is likely to deploy on AMD and Nvidia first and then transition to in-house silicon later.

It’s also worth noting that AMD’s MI455X and Nvidia’s Rubin GPUs, also expected to ramp production in early 2027, are very different kinds of chips optimized for a mix of training and inference, whereas OpenAI’s custom silicon only needs to excel at one job: inference.

When it comes to inference, compute is key but memory bandwidth is king.

And based on early benchmarks OpenAI shared with the press before its Hot Chips presentation Tuesday, the chip is shaping up to be an inference beast.

Testing on SemiAnalysis’ InferenceX benchmark suite — presumably this is an unofficial test — shows OpenAI’s Jalapeño-based systems delivering between 1.5x and 1.9x more “AI work” at peak throughput, and 1.7x to 3.6x lower end-to-end latency than the competition across GPT-OSS-120B, DeepSeek R1, and Kimi K2.5.

If the latter two seem like weird models for OpenAI to be testing against, it's not that OpenAI plans to use these chips to run competitors' models, it’s just the models InferenceX uses. In any case, the test shows that Jalapeño isn’t some model-specific architecture designed for maximum performance at the expense of programmability… cough, cough Taalas.

Meanwhile, for ultra-low-latency inference, which has become the hot new segment for AI infrastructure providers, OpenAI says its chips are 2.1x to 4.1x faster.

At a system level — we're starting here because the frontier models OpenAI trains rarely run on a single chip any more — each Jalapeño system with its 128 accelerators packs 1.7 exaFLOPS of 4-bit compute, 27.5 TB of HBM4 and just shy of 2 petabytes a second of memory bandwidth.

By comparison, AMD and Nvidia’s latest rack systems are faster, delivering 1.46x to 2x more compute and up to 12 percent more memory on Helios, but just 85 percent the memory bandwidth of OpenAI’s rack.

As of writing, OpenAI hasn’t shared system-level power consumption, but based on what we know about the accelerators we’d wager each rack will use between 40 and 60 percent of the power of competing GPU systems.

Source: Read the original article on www.theregister.com