OpenAI told Hot Chips that Jalapeño will trickle out this year and reach volume in 2027. SemiAnalysis measured 1.5x to 1.9x more inference work on selected models using A0 samples still in the lab. A chip can beat Blackwell on a bench and still miss Nvidia's share this year.
OpenAI used Tuesday’s Hot Chips session at Stanford to put numbers on Jalapeño, the Broadcom-built inference ASIC it unveiled in June. Hardware VP Richard Ho said racks of 128 chips will deliver 1.7 exaFLOPS of 4-bit compute and 27.5 terabytes of HBM4, at a 700-watt TDP that held at or below 550 watts on the workloads tested. SemiAnalysis, invited into OpenAI’s lab, ran its InferenceX suite on A0 silicon and reported 1.5x to 1.9x more “AI work” at peak throughput, and 1.7x to 3.6x lower end-to-end latency, versus competing GPU systems on GPT-OSS-120B, DeepSeek R1, and Kimi K2.5.
The dominant reading writes itself. A custom chip, nine months from design to tape-out, beats Blackwell on inference, Broadcom collects another XPU, and Nvidia’s share is next. Hock Tan already promised gigawatt-scale data centers with Microsoft “beginning in 2026.” That is the hedge-against-Nvidia slide. It is not the residual.
A Faster Sample Is Not a Fleet
If Jalapeño already outperforms Blackwell on the reported inference workloads, why is OpenAI’s deployment expected to remain small-volume through the end of 2026?
The first constraint is manufacturing, not the scoreboard. SemiAnalysis recorded those wins on the A0 stepping. B0 — the die quoted at 13.4 petaFLOPS of MXFP4 on TSMC N3P, with a claimed 25 percent perf-per-watt lift — is still in the fab. Production, the same report says, ramps gradually over 2027, with most output scheduled for the end of next year. Ho told DCD the chip appears in limited quantities later this year and more widely next year. A full pod is 2,048 ASICs. You do not fill a gigawatt from a characterization board.

The second constraint is what the bench did not measure. SemiAnalysis verified InferenceX in person and did not run its AgentX suite — the long-context, multi-turn load that stresses routers and prefix cache the way production agents do. The models are not the open frontier. The comparison set is still heavy on GB200 and GB300; Vera Rubin systems are shipping now, on the same HBM4 generation Jalapeño needs. Jalapeño’s published points omit speculative decoding, which SemiAnalysis notes can add a 3x-plus cost-per-token swing on GPU stacks. A selected-workload lead can survive all of that and still leave most of OpenAI’s serving mix on CUDA.
The third constraint is the job the chip refuses. Jalapeño is inference-only. Training, and the first deploy of new models, still wants a programmable GPU. Ho was explicit: the strategy includes “very, very good partners at Nvidia, Cerebras, and others,” and OpenAI will not meet its compute needs “with one shot.” It will not sell the chip. It is “struggling to have enough.” That is the opposite of a 2026 share raid. It is a tokens-per-megawatt project inside a power-limited data-center envelope — the same envelope Jensen has been pricing as revenue per watt, which is why Nvidia’s record quarter still got shrugged and why Groq-style inference capacity remains a procurement object, not a die.
Markets will still misprice the Hot Chips graph as a 2026 TAM haircut. Earnings-week Nvidia is a shipping company with a software stack. Jalapeño is a lab-proven ASIC whose TCO story, even in SemiAnalysis’s own framing, is head-to-head with Rubin on tokens per dollar until speculative decoding lands. Utility still has to fill the racks. A cheaper token is not a trained model.
Small volume through 2026 is the expected output of A0-in-lab, B0-in-fab, inference-only silicon running 8k1k benches — not a contradiction of the benchmark.
The test that would kill this reading is simple. Documented, large-scale Jalapeño pods in production this year, or an AgentX-class result showing the A0 advantage was an artifact that cannot ship. Until then, watch TSMC N3P and CoWoS allocations into 2027, OpenAI’s next Nvidia and AMD purchase disclosures, and whether Ho’s “limited quantities” become a 10-Q line. The chip can be real. The share shift can still wait.
Continue reading
Sources
OpenAI Hot Chips 2026 remarks via The Register and Data Centre Dynamics; SemiAnalysis Jalapeño lab/InferenceX report (Aug. 25, 2026); OpenAI-Broadcom June 24, 2026 product announcement; Richard Ho presentation notes