OpenAI has fully unveiled Jalapeño, the company's debut AI accelerator chip. The chip is designed to deliver up to 13.4 petaflops of 4-bit compute and features 232 gigabytes of high-bandwidth memory (HBM4), with a memory bandwidth of 15.4 terabytes per second. According to benchmarks cited by OpenAI, Jalapeño can reduce end-to-end latency by up to 3.6 times compared to Nvidia’s GB300 while consuming less power.
The development of Jalapeño was significantly accelerated by OpenAI's large language models (LLMs). The process moved from the initial architecture concept to the first silicon in under 20 months, with only nine months between the first RTL (register-transfer level) code and the tape-out. OpenAI’s hardware vice president, Richard Ho, stated that the models allow engineers to explore more design paths much faster, although humans remain the final decision-makers.
The design workflow utilized Accelerated Hardware Synthesis (XLS), an open-source high-level synthesis tool, which allows engineers to write designs in programming languages like C++ or DSLX. These are then converted into Verilog, a hardware description language. The team also applied internal AI models to optimize software for the hardware. In one instance, internal AI-driven software optimization improved performance on a DeepSeek benchmark from 0.31 percent to 88.94 percent of the theoretical ceiling within approximately 40 hours.
While OpenAI handled the end-to-end system design, including the inference accelerator and memory hierarchy, it partnered with Broadcom to manage physical design from the gates onward. This collaboration helped meet the rapid development timeline. OpenAI plans to deploy Jalapeño in pods containing 2,048 chips and aims to integrate the lessons learned from this project into future commercial LLMs.
Sources
- How OpenAI Used Its Own LLMs to Design Its Jalapeño Chip (Hacker News Frontpage, 2026-09-18)