OpenAI and Broadcom announced Jalapeño on Tuesday, OpenAI’s first custom AI accelerator and the opening move in a multi-generation co-design partnership the two companies have been working on quietly for the last year. Engineering samples are running in the lab at production target frequency and power. The workloads being tested include GPT-5.3-Codex-Spark, which is the production model OpenAI uses to write code, including, by Broadcom’s own description, large portions of the verification and optimization stack used to build Jalapeño itself. The chip helped design the next version of itself. We have entered the part of the cycle where the joke writes itself.

The headline economic number is roughly 50 percent cost savings per inference compared to current state-of-the-art AI GPUs, which in 2026 means Nvidia Blackwell and Rubin parts. That figure is first-party, lab-stage, and should be treated with the usual skepticism that applies to vendor benchmarks, but the underlying architecture story holds up independent of the marketing. Jalapeño is a blank-slate design targeting LLM inference specifically, rather than a general-purpose accelerator retrofitted from training silicon. That means narrower memory hierarchies tuned for attention patterns, more aggressive sparsity handling, and on-die data movement optimized for the actual shape of GPT-5-class workloads rather than the average of every model architecture from 2018 onward. If even half of the 50 percent claim survives contact with production scale, the unit economics of running ChatGPT change materially.

Broadcom’s role here is the more interesting commercial story. The press release says the silicon went from initial design to manufacturing tape-out in nine months, which both companies describe as one of the fastest cycles ever achieved for an advanced-node ASIC of this complexity. Yahoo Finance reported separately that Broadcom built the chip in record time but that the financial upside of the program largely accrues to OpenAI rather than Broadcom, because OpenAI controls the IP and the deployment relationship with Microsoft for the gigawatt-scale data centers that will host it starting later this year. Broadcom’s compensation is high-margin design-services revenue plus the next several generations of follow-on tape-outs. OpenAI is buying its way off the Nvidia roadmap one custom chip at a time, and Broadcom is the courier.

The strategic read is that the inference market is splitting into two: hyperscalers running custom silicon designed against their own workloads, and everyone else still buying merchant GPUs. Google has TPUs, Amazon has Trainium and Inferentia, Anthropic just signed a multi-gigawatt deal for next-gen TPUs with Google and Broadcom in April, and OpenAI now has Jalapeño with more generations promised. Nvidia is still the only place to go if you do not have an in-house frontier model to justify the fixed cost of a custom silicon program, which describes essentially every AI customer that is not a top-five lab. Jensen’s keynote at the next GTC is going to be very interesting.

openaibroadcomjalapenocustom-siliconasicinference-chipllm-acceleratorgpt-5microsoftgigawatt-buildout