The HBM Gamble: Can OpenAI's Jalapeño Break the Inference Bottleneck?

AI-generated image · US National Wire
OpenAI is betting on a custom silicon beast to outpace Nvidia and AMD, but the success of the Jalapeño rack hinges on whether its massive HBM4 memory footprint is sustainable.
Sam Altman and the team at OpenAI are attempting a high-stakes pivot toward custom silicon to solve the industry's most pressing constraint: inference latency. As The Register first reported, OpenAI recently provided a first look at its Jalapeño AI accelerator during the Hot Chips semiconductor development conference at Stanford. Developed in collaboration with Broadcom, Jalapeño represents the first in a series of custom chips designed specifically for inference, rather than the general-purpose training and inference mix found in competing hardware.
From a hardware nerd's perspective, the Jalapeño system is an absolute monster. Each rack consists of 128 accelerators, providing a combined total of 1.7 exaFLOPS of 4-bit compute. However, the real story isn't the raw compute—it's the memory. Each chip is equipped with 216 GB of HBM4 memory, which The Register notes is likely six 12-high stacks. At a system level, this totals 27.5 TB of HBM4, providing just under 2 petabytes per second of memory bandwidth.
OpenAI is positioning this as a direct strike against the 'Blackwell bottleneck.' According to early benchmarks from SemiAnalysis’ InferenceX suite cited by The Register, Jalapeño systems delivered 1.5x to 1.9x more 'AI work' at peak throughput and 1.7x to 3.6x lower end-to-end latency compared to the competition across models like DeepSeek R1, Kimi K2.5, and GPT-OSS-120B. For ultra-low-latency workloads, OpenAI claims its chips are 2.1x to 4.1x faster than the Nvidia GB200 NVL72 and GB300 NVL72 rack systems.
**Leo's Take: The Yield Risk**
Opinion: While the performance numbers are seductive, the viability of Jalapeño rests entirely on the yield and cost of that HBM4. Packing 27.5 TB of high-bandwidth memory into a single rack is a massive engineering lift. If the HBM4 yields are low, the cost per rack could skyrocket, turning a performance win into a financial liability. OpenAI is betting that the efficiency gains—including a claim that racks will use 40% to 60% of the power of competing systems—will offset the silicon risk.
To achieve these optimizations, Richard Ho, VP of hardware at OpenAI, noted that the chip's architecture focuses on reducing data movement. By keeping model state and KV caches local, the system avoids the communication delays that plague larger clusters. This is supported by what appears to be a large SRAM cache, allowing the chip to handle both compute-heavy prefill and memory-intensive decode phases efficiently.
Despite the ambition, OpenAI isn't cutting ties with its investors. The Register reports that OpenAI will likely continue deploying on Nvidia and AMD—specifically citing the MI455X and Rubin GPUs—for training tasks, as those programmable GPUs remain superior for that specific workload. Jalapeño is a specialist tool, and its success depends on whether OpenAI can actually move these into volume production by 2027 without the HBM costs breaking the bank.

