OpenAI Posts First Jalapeño Inference Chip Benchmarks at Hot Chips
At Hot Chips 2026, OpenAI published its first benchmark disclosure for Jalapeño, a custom inference ASIC co-developed with Broadcom and sized at roughly 700W against NVIDIA's 1,400W flagship GPUs. In OpenAI's internal tests on real model workloads, the company reported 1.5 to 1.9 times more work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency versus GB200 and GB300 systems, with production deployment into its own infrastructure expected by year-end. Coverage from Bloomberg, CNBC, ServeTheHome, and Tom's Hardware frames the reveal as hyperscaler vertical integration that could pressure NVIDIA's economics, though several linked articles offered only headlines or paywall boilerplate. Operators should treat the efficiency and latency figures as vendor-reported until independent verification on disclosed workloads and methodology arrives.
OpenAI Posts First Jalapeño Inference Chip Benchmarks at Hot Chips
OpenAI released first benchmark details for its custom inference chip Jalapeño, claiming materially better efficiency and latency than NVIDIA GB200/GB300 systems on real model workloads. In OpenAI’s tests, Jalapeño delivered 1.5–1.9× more work per watt at peak throughput and 1.7–3.6× lower end-to-end latency, with deployment into OpenAI infrastructure slated to begin by year-end.
Key takeaway
OpenAI's Hot Chips Jalapeño benchmarks mark a credible move from rumor to disclosed silicon, but the headline efficiency wins remain self-reported pending third-party replication.
What happened
OpenAI disclosed first benchmark results for its Jalapeño custom inference chip at Hot Chips 2026, according to Latent Space AINews coverage and corroborating trade press. The company claimed materially superior efficiency and latency compared with NVIDIA GB200/GB300 systems on production-relevant model workloads.
Reported figures include 1.5–1.9× more work per watt at peak throughput and 1.7–3.6× lower end-to-end latency in OpenAI's tests, with Tom's Hardware noting a 700W Broadcom co-developed ASIC benchmarked against a 1,400W NVIDIA flagship. OpenAI indicated deployment into its own infrastructure should begin by year-end.
Evidence
OpenAI released first Jalapeño benchmark details claiming better efficiency and latency than NVIDIA GB200/GB300 on real model workloads.
Latent Space · attributed
OpenAI released first benchmark details for its custom inference chip Jalapeño, claiming materially better efficiency and latency than NVIDIA GB200/GB300 systems on real model workloads.
OpenAI's tests showed 1.5–1.9× more work per watt and 1.7–3.6× lower end-to-end latency, with year-end deployment planned.
Latent Space · attributed
In OpenAI's tests, Jalapeño delivered 1.5–1.9× more work per watt at peak throughput and 1.7–3.6× lower end-to-end latency, with deployment into OpenAI infrastructure slated to begin by year-end.
Tom's Hardware reported a 700W Jalapeño ASIC co-developed with Broadcom claiming up to 1.9× throughput per kilowatt and 3.6× lower latency versus a 1,400W NVIDIA flagship GPU.
Tom's Hardware AI · attributed
OpenAI's 700W Jalapeño ASIC outpaces 1,400W Nvidia flagship GPU — claims up to 1.9x throughput per kilowatt and 3.6x lower latency, co-developed with Broadcom
Bloomberg reported OpenAI said its Jalapeno chips performed better than NVIDIA's current lineup.
Bloomberg Technology · attributed
OpenAI says its new Jalapeno chips performed better than Nvidia's current lineup.
ServeTheHome reported OpenAI showcased a custom AI ASIC codenamed Jalapeno at Hot Chips 2026.
ServeTheHome AI · attributed
ServeTheHome reported that OpenAI is showcasing a custom AI ASIC, reportedly codenamed 'Jalapeno', at Hot Chips 2026.
CNBC framed OpenAI's Jalapeño custom silicon as a new threat to NVIDIA margins amid broader hyperscaler accelerator roadmaps.
CNBC AI · attributed
OpenAI's Jalapeño AI chip brings new 'threat' to Nvidia margins as custom silicon gains ground
Why it matters
Frontier labs designing bespoke inference ASICs shifts competitive advantage from GPU purchases toward vertically integrated serving economics and software stacks tuned to proprietary hardware.
Limits and uncertainties
Efficiency and latency figures come from OpenAI's own tests; several trade-press excerpts in the packet contain headlines or navigation UI rather than methodology or independent verification.
Stratechery and The Information items in the packet surfaced only paywall or language-selector boilerplate, so their analytical claims could not be substantiated from supplied text.
The Cerebras Hot Chips item paired a technical headline with SEC risk-factor boilerplate, offering no usable technical benchmark signal in the packet.
Practical implications
Operators should prioritize hardware-abstraction layers and portable inference runtimes so workloads are not locked to a single accelerator vendor.
Builders should assume OpenAI may increasingly optimize serving stacks for Jalapeño-class silicon rather than generic GPU paths.
Capacity planners should treat year-end deployment as a watchpoint, not a confirmed production milestone, until rollout evidence appears.
What to watch
Whether OpenAI begins deploying Jalapeño into its own infrastructure by year-end as stated.
Independent replication of OpenAI's GB200/GB300 comparison benchmarks on named workloads.
NVIDIA earnings guidance and any public response to custom-silicon efficiency claims.