DistantNews
Support us
OpenAI’s First ASIC, ‘Jalapeño,’ Was Built to Make the Chips It Needs, Not Beat Nvidia

OpenAI’s First ASIC, ‘Jalapeño,’ Was Built to Make the Chips It Needs, Not Beat Nvidia

From Dong-A Ilbo · () Korean

Translated from Korean and summarized by DistantNews. Read the original for the full story.

At a glance

News Documents & data New plan
  • OpenAI presented Jalapeño, a custom ASIC developed with Broadcom for large-language-model inference, at Hot Chips 2026.
  • The chip uses 216 GiB of HBM4 memory with 15.4 TB per second of bandwidth and contains 64 core and memory slices.
  • OpenAI said its design targets lower latency and energy use in inference rather than directly competing with Nvidia across all workloads.

OpenAI used an unusually casual presentation title at the formal Hot Chips 2026 semiconductor conference: “You Can Just Build Things … Chips.” The phrase captured the company’s approach as it introduced Jalapeño, a custom accelerator designed for large-language-model inference.

OpenAI first disclosed the chip with Broadcom in June, but initially released no detailed performance figures, benchmarks, pricing, deployment scale or power-efficiency data. At the conference, it provided a fuller account of the design and development process.

The company said it began work on the architecture concept in October 2024, started the initial RTL implementation in February 2025 and completed the logic design five months later. It taped out the design in November 2025 and produced its first chip in May 2026. OpenAI therefore described nine months from internal design to tape-out, while the first silicon took 19 months from the start of architecture work. The company’s nine-month figure does not include the earlier architecture-definition period.

Jalapeño has a central compute die flanked by three HBM4 stacks on each side. The memory offers 216 GiB of capacity and 15.4 TB per second of bandwidth. The chip contains 64 core slices matched with 64 HBM slices, allowing frequently used weights and computation data to stay near specific cores. OpenAI said this reduces waiting time and avoids relying on a single shared-memory pathway for every operation.

The company reported 13.4 petaflops for low-precision MXFP4 matrix operations. In a comparison using the GPT-OSS 120B model and SemiAnalysis’s InferenceX benchmark, OpenAI said Jalapeño generated up to 1.9 times more tokens per kilowatt than Nvidia’s GB200. Its stated goal, however, is targeted inference performance, low latency and lower energy use, not simply defeating Nvidia.

You Can Just Build Things … Chips

— OpenAITitle of OpenAI’s presentation session at Hot Chips 2026.
About this summary

Originally published by Dong-A Ilbo in Korean. Translated, summarized, and contextualized automatically by DistantNews, with a note on how the source frames the story. Not individually reviewed before publishing. How this works.