OpenAI unveils new details of AI accelerator jalapeno pepperDeveloped in partnership with Broadcom, writes register. how assertion OpenAI, Jalapeño designed for inference, showing better results in tests than current NVIDIA family of accelerators – GB 200 and GB 300.
SemiAnalysis InferenceX benchmark suite tests show that the Jalapeño-based OpenAI system improves AI workloads at peak throughput by 1.5 to 1.9 times and achieves end-to-end latency of 1.6 times compared to competitors on GPT-OSS-120B, DeepSeek R1 670B, and Kimi K2.5 1T.
In internal testing, Jalapeño also performed well on some larger, more advanced OpenAI models that have not yet been released. This shows that chip design “becomes more valuable as workloads grow and become more complex,” the company said in a blog post.
At the system level, a rack with 128 Jalapeño accelerators delivers 1.7 EFLOPS of 4-bit compute with 27.5 TB of HBM4 for a total throughput of just under 2 PB/s. The chip combines a compute die with six HBM4 stacks totaling 216 GiB (15.4 TB/s), delivering 13.4 PFLOPS of MXFP4 performance and scalable to 27 EFLOPS and 432 TiB of memory in a 2,048 ASIC system.

Image source: OpenAI/ServeTheHome
The agency said that because the chip performs well at low power consumption of 700 watts, Jalapeno will allow OpenAI to save money when operating its data centers, and power is a key cost item. Bloomberg Richard Ho, Vice President of Hardware at OpenAI. The latest AMD and NVIDIA rack-mount systems are faster, offering 1.46x to 2x more computing power and up to 12% more memory (Helios), but only 85% of the memory bandwidth of the OpenAI rack.
Richard Ho said the chip is designed to minimize data movement and intermediate storage. “We designed Jalapeño to minimize data movement and communication latency. This means that model state (including the KV cache used to generate responses) can be explicitly allocated and stored locally, while the system activates the right combination of compute, memory and network resources for each inference stage.” the company said in a blog post.
Unlike NVIDIA racks Gronk LPXJalapeños are not a be-all and end-all solution. The company told The Register that it was optimized for the resource-intensive prefill operation of processing prompts and the memory-intensive decoding stage of generating output tokens.
It is also worth noting that by using artificial intelligence to design the architecture and optimize the chip’s reasoning capabilities, OpenAI took only nine months from the initial design to the development of the first proof-of-concept chip. The process involves using custom models to optimize the inference engine and writing custom cores when new models are introduced.
Jalapeño’s speed advantage is so great that OpenAI can achieve results that were previously only possible by using different types of memory and designed chips. For example, currently OpenAI uses chips brain For some of its LL.M.s, Ho said smaller models are best suited. Jalapeño, on the other hand, can handle larger models, he noted.
However, the company has no plans to abandon cooperation with other suppliers, including NVIDIA. “We have a huge need for computing power, which is why we attract so many different suppliers, – he said. — This will last for a while. ” The jalapeños are expected to be released in limited quantities later this year, with production ramping up further next year.
Ho said a second version of Jalapeño is in the advanced stages of development and is expected to go into production “in the coming months.” Concepts for third-generation chips are also already under development. The Register urges caution with OpenAI’s claims. The new chips were tested against, but not compatible with, NVIDIA GB200 NVL72 and GB300 NVL72 systems due to be released in 2024 and 2025 respectively. Vera Rubindelivery has just begun.
Source:
