OpenAI Built Its Own Chip. It's Called Jalapeño. It's Fast.

OpenAI Built Its Own Chip. It's Called Jalapeño. It's Fast.

Every AI lab eventually hits the same midlife crisis: buy more Nvidia GPUs, or build your own silicon and pretend you always planned to. OpenAI just made its choice, and it named the result after a pepper — presumably because "chip that makes your GPU bill sweat" didn't fit on a slide.

First-Party Silicon, First Real Numbers

OpenAI unveiled Jalapeño, its first custom inference chip, developed with Broadcom, and this week published its first head-to-head results rather than just a roadmap slide. Tested on the public InferenceX benchmark across GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T, Jalapeño delivered 1.5x to 1.9x more peak AI throughput per watt and 1.7x to 3.6x lower end-to-end latency than the commercial systems it was benchmarked against.

For the highly interactive workloads that actually annoy users — the ones where a chatbot pauses just long enough to feel broken — OpenAI measured 2.1x to 4.1x higher performance. The company says it plans to start deploying Jalapeño inside its own infrastructure by the end of 2026, while still buying Nvidia and other accelerators alongside it. This is Gen 1; Gen 2 is already deep in development.

Why Everyone's Suddenly a Chip Company

Custom silicon used to be the kind of multi-billion-dollar, decade-long bet only Google or Amazon would make. Now it's table stakes for any AI company that wants to control its own margins instead of watching Nvidia's earnings calls with a knot in its stomach. OpenAI joins Google's TPUs and Amazon's Trainium in the "yes, we make our own chips now" club.

The part worth watching isn't the benchmark chart — vendor-run comparisons always flatter the vendor — it's the strategic shift underneath it: when a model company also designs the memory, networking, and silicon it runs on, it starts controlling the entire cost curve of inference, not just the software layer. That's the actual moat forming here.

Whether Jalapeño out-cooks Nvidia in production remains to be seen, but the era of "just rent GPUs and worry about it later" is officially over.

Faster, cheaper inference eventually trickles down to every AI feature businesses build on top of it — if you're planning how AI fits into your own product roadmap, WTK's a good sounding board — get in touch.

Source: OpenAI