US National WireUS NATIONAL WIRE
Tech

OpenAI's Jalapeño Chip Targets Inference Efficiency to Break GPU Dependency

Portrait of Leo Abernathy
Leo Abernathychips & semiconductorsAug 25AI
OpenAI's Jalapeño Chip Targets Inference Efficiency to Break GPU Dependency

AI-generated image · US National Wire

Developed with Broadcom, the custom silicon prioritizes memory bandwidth and low latency over the general-purpose flexibility of Nvidia and AMD hardware.

OpenAI is attempting to shift its hardware dependency by developing Jalapeño, a custom AI accelerator designed specifically for inference. Unveiled at the Hot Chips conference, the chip was created in collaboration with Broadcom and utilized OpenAI's own models during its development, according to TechCrunch and The Register.

While Nvidia's Rubin and AMD's MI455X GPUs are optimized for a blend of training and inference, Jalapeño focuses exclusively on the latter. This specialized approach allows OpenAI to prioritize memory bandwidth—a critical factor for inference performance. A Jalapeño system featuring 128 accelerators delivers 1.7 exaFLOPS of 4-bit compute and 27.5 TB of HBM4. The Register reports that while competing rack systems from Nvidia and AMD may offer more raw compute, OpenAI's system provides superior memory bandwidth, reaching nearly 2 petabytes per second.

According to Richard Ho, OpenAI's VP of hardware, the architecture is designed to minimize data movement and communication delays. By keeping model state and KV caches local, the chip optimizes both compute-heavy prefill operations and memory-intensive decode phases. TechCrunch notes that this full-stack approach allows OpenAI to address specific bottlenecks in the inference process.

Early benchmarks using SemiAnalysis’ InferenceX suite indicate that Jalapeño-based systems provide 1.5x to 1.9x more throughput and 1.7x to 3.6x lower end-to-end latency compared to the Nvidia GB200 NVL72 and GB300 NVL72 racks. Richard Ho stated that the chip can serve more AI work per unit of power while returning responses faster.

OpenAI expects small volumes of the chip to deploy by the end of 2026, with volume production ramping up in 2027. However, The Register notes that OpenAI will likely continue using Nvidia and AMD hardware for training tasks due to the programmable nature of GPUs.

Sources

More from Leo Abernathy