OpenAI's Jalapeño chip posts its first numbers — and it still can't escape the memory squeeze
TL;DR: On 25 August 2026 OpenAI published the first benchmark results for Jalapeño, the custom inference chip it announced with Broadcom in June. The claims: 1.5-1.9x more throughput per kilowatt and 1.7-3.6x lower end-to-end latency than Nvidia GB200 and GB300 systems, from a 700W package competing against 1,200-1,400W parts. SemiAnalysis ran its public InferenceX suite at OpenAI’s labs and confirmed the chip is real — while stating that OpenAI supplied every number. For you: nothing. The chip is not for sale, serves almost no traffic until 2027, and was measured on open-weight models rather than the ones you pay for. The part that should update your model is narrower and sharper: Jalapeño cuts out Nvidia’s margin while still buying six stacks of HBM4, the exact component that just pushed AI server prices up more than 15%. Custom silicon routes around the vendor, not the shortage.
What was actually published
| Metric | Jalapeño vs GB200 / GB300 |
|---|---|
| Throughput per kilowatt | 1.5-1.9x higher |
| End-to-end latency | 1.7-3.6x lower |
| Interactive workloads | 2.1-4.1x better |
| Package power | 700W vs 1,200-1,400W |
| Memory | 6x HBM4 stacks — 216 GiB, 15.4 TB/s |
| Process / partners | TSMC 3nm, Broadcom ASIC, Samsung HBM4 (reported) |
| Benchmark | SemiAnalysis InferenceX (Apache 2.0) |
| Models tested | GPT-OSS 120B, DeepSeek R1 670B, Kimi K2.5 1T |
| Deployment | Very small volumes end 2026; scale 2027 |
| Availability | Internal to OpenAI. Not sold, not rentable |
The single most striking figure is on user-perceived speed: roughly 700 tokens per second per user on Kimi K2.5, which SemiAnalysis put at about 9x the next best chip it has measured. On GPT-OSS, Jalapeño reached nearly double GB200’s highest throughput point. Sam Altman’s summary was four words: “we made a chip and it is fast.”
Read the asterisks
The framing across most coverage — OpenAI beats Nvidia — is doing more work than the evidence supports. SemiAnalysis, whose benchmark this is, was notably careful, and the caveats it published are the story:
- “All numbers are provided to us by OpenAI.” SemiAnalysis visited the labs and confirmed the runs happen. It did not independently operate the hardware.
- The full suite was not run, and there are no AgentX results — the agentic workloads that stress routers and prefix caching, which is what production traffic actually looks like.
- Only A0 silicon was tested. A B0 stepping already exists and was not measured.
- The models are not frontier. GPT-OSS, R1 and K2.5 are open-weight checkpoints. Nothing here shows how Jalapeño serves GPT-5.6.
One caveat cuts the other way, and fairness requires stating it: Jalapeño ran single-token prediction with no speculative decoding, while the Nvidia comparison systems ran optimised multi-token prediction. On a like-for-like basis the gap would likely widen, not narrow.
So this is not a fabrication, and dismissing it as marketing would be lazy. It is a well-instrumented demo with an independent witness — meaningfully better than a vendor slide, meaningfully worse than a measurement anyone can reproduce. Nobody outside OpenAI has hands on this hardware, and until someone does, the numbers are a claim with good manners.
The point most coverage missed
Set this against yesterday’s memory story, because the two together say something neither says alone.
There are two separate costs inflating AI inference, and they are constantly conflated:
- Nvidia’s margin — gross margins in the 70-75% range, paid by everyone renting Blackwell-class capacity.
- The memory bill — DRAM and HBM4 repricing, the stated driver of the 15%+ increase on systems shipping in early 2027, and a shortage that Gartner expects to run through at least H1 2027.
Jalapeño attacks the first and is fully exposed to the second. Each package reportedly carries six HBM4 stacks. OpenAI is not escaping the memory market; it is bidding into it, and its 10GW Broadcom commitment makes it one of the largest new sources of HBM4 demand in the industry — plausibly tightening the shortage it is subject to.
That reframes the “threat to Nvidia margins” headline into something more precise. Custom silicon lets a hyperscaler stop paying an intermediary’s markup. It does not manufacture memory. The floor under how cheap inference can get is set by the shortage, not by Nvidia’s price list, and Jalapeño does not move that floor for anyone — including OpenAI.
Everyone has a hedge now
The genuinely durable development is not one chip. It is that the vertical-integration race is effectively over as a question of whether:
| Lab | Silicon position |
|---|---|
| TPU line, now architecturally co-designed with Gemini — and rentable | |
| Anthropic | Multi-sourced across AMD MI450 at 2GW and TPU capacity |
| OpenAI | Jalapeño gen-1 deploying; gen-2 near tape-out, gen-3 in design |
| Amazon / Microsoft | Trainium and Maia, both merchant-adjacent |
For a buyer, the consequence is not about any single vendor’s cost curve. It is that the scenario people reasonably worried about in 2024 — a single GPU supplier’s pricing squeezes every model vendor at once, and API prices rise in lockstep — is now much less likely. The labs have different cost structures, different suppliers and different exposure, several of them backed by very large dedicated financing structures. Divergent cost structures sustain price competition. That is a mild but real argument that the price war persists rather than resolving into comfortable oligopoly pricing.
What does not change
Being blunt about the null result, because it is most of the practical answer:
- Your API bill. Jalapeño serves negligible traffic in 2026. The rates that matter are the ones with clocks already on them — Sol’s promotional $4/$20 expiring 21 November, Gemini 3.7 Flash’s introductory rate to 31 December.
- Your infrastructure options. You cannot buy, rent or target this chip. Unlike TPUs, it is not a product.
- Self-hosting economics. If anything these get slightly worse, since the real hardware cost of running frontier open weights is dominated by memory capacity — the constrained input, not the one Jalapeño addresses.
- Model quality. A faster serving substrate does not make a model smarter. It buys latency and headroom, which show up as fewer capacity-driven degradations rather than better answers.
What to do
- Do not reprice your 2027 budget on this. Post-promotional rates remain the correct planning assumption. A chip advantage inside one vendor is not a committed discount to you.
- Watch latency SLAs, not price lists. If Jalapeño delivers anything user-visible in 2027, the first place it appears is time-to-first-token and sustained throughput under load on OpenAI endpoints — a real advantage for voice, agents and interactive coding. Instrument that now so you can detect it.
- Keep switching cheap. Divergent silicon strategies mean the price and latency leader will keep changing hands. Keep ChatGPT, Claude and Gemini evaluated in parallel and treat selection as routing. For coding specifically, the best AI coding tools guide tracks which harnesses are model-agnostic — that flexibility is the asset here.
- Discount vendor-supplied benchmarks by default. Not to zero. But the reproducibility standard you would apply to a model benchmark applies to silicon too, and this one does not meet it yet.
- Track HBM4, not chip announcements. Memory supply is the variable that actually governs inference cost and long-context availability through 2027. It is also the one nobody issues a press release about.
The honest uncertainty
The strongest counter-argument is timing, and it is not weak. Jalapeño gen-1 is being compared against Blackwell-generation systems while Nvidia’s Vera Rubin generation is already reaching customers. By the time Jalapeño carries serious volume in 2027, the honest comparison set will have moved. OpenAI’s answer is the roadmap — gen-2 near tape-out, gen-3 in design, on a nine-month design-to-tape-out cadence that is genuinely unusual — but a roadmap is a promise and Vera Rubin is a shipping product.
The claim that survives all of that is small and worth holding onto: an inference ASIC built by a model company, benchmarked on end-to-end serving rather than peak FLOPs, is now credible enough that the burden of proof has shifted. A year ago the question was whether anyone outside Google could make custom inference silicon work. That question is closed. The open one is whether it arrives fast enough to matter before the memory shortage sets the price anyway.
Update, 27 August 2026 — Nvidia’s answer to the custom-silicon squeeze appears to be distribution. A day after these benchmarks landed, The Information reported that Nvidia has agreed to acquire Hugging Face for $12.9 billion (Business Insider reports unresolved talks above $13 billion; neither company has confirmed). Read alongside this article the logic is coherent. Nvidia cannot buy back the demand that closed labs are moving onto their own inference chips — those customers are leaving deliberately. What it can defend is the open-weight remainder, and that segment reaches buyers through one hub. Paying roughly 86 times revenue for a business earning about $150 million a year is not a software multiple; it is the cost of not being disintermediated on the one flank still open. The full read, including what buyers should check this week.
Frequently asked questions
Can we buy or rent Jalapeño capacity?
No, and there is no announced plan to change that. OpenAI is deploying Jalapeño inside its own infrastructure, not selling it as merchant silicon and not exposing it as a distinct rentable instance type the way AWS sells Trainium or Google sells TPU access. The deployment schedule is also slower than the headlines imply — very small volumes at the end of 2026, with meaningful capacity only arriving through 2027. The only way you will ever touch Jalapeño is indirectly, by calling OpenAI's API and being served by whatever hardware happens to be behind it, which OpenAI does not disclose per request. That is a genuine difference from Google, where TPUs are both an internal cost advantage and a product you can rent. For procurement purposes, treat Jalapeño as a fact about OpenAI's future cost structure rather than an option in your own infrastructure plan.
Are the benchmark numbers trustworthy?
They are credible in direction and unverifiable in magnitude, and the distinction matters. The positives are real: InferenceX is a public, Apache 2.0-licensed benchmark suite from SemiAnalysis rather than a private vendor harness, SemiAnalysis was invited to OpenAI's labs and confirmed in person that the chip and the benchmark runs are real, and measuring end-to-end request serving is a much harder test to game than peak FLOPs. The limits are equally real, and SemiAnalysis states them plainly: every number was provided by OpenAI, the full InferenceX suite was not run, no AgentX results exist, and only A0 silicon was tested when a B0 stepping already exists. Two methodology details deserve particular weight. Jalapeño ran single-token prediction with no speculative decoding while the Nvidia comparison systems used optimised multi-token prediction, which if anything understates Jalapeño's ceiling. And the models tested — GPT-OSS 120B, DeepSeek R1 670B, Kimi K2.5 1T — are open-weight rather than frontier, so none of this demonstrates that the chip serves GPT-5.6 well. The honest summary is a well-instrumented demo with independent supervision, not an independent measurement.
Does this mean OpenAI's API prices will fall?
Not on this news, and not on this timeline. Nothing about Jalapeño changes OpenAI's cost per token in 2026, because the chip serves almost no traffic this year. Even in 2027, a cost advantage on a portion of the inference fleet does not mechanically become a price cut — frontier API rates are set to win market share, not on cost-plus, which is why several published rates are widely believed to sit at or below fully-loaded serving cost. What Jalapeño changes is capacity to sustain aggressive pricing rather than the immediate incentive to cut. The concrete near-term prices remain the ones already on the calendar: GPT-5.6 Sol's promotional $4/$20 expires on 21 November 2026, and Gemini 3.7 Flash's introductory rate runs to 31 December. Budget against the post-promotional rates, not against a chip that will not be serving your traffic at scale until well after those clocks run out.
Why does the memory point matter more than the Nvidia comparison?
Because it separates the two different costs people keep merging. There are two things making AI inference expensive right now: Nvidia's gross margin, which has run in the neighbourhood of 70-75%, and the price of high-bandwidth memory, which repriced sharply through 2026 and is the stated driver of the more than 15% increase on AI servers shipping in early 2027. Custom silicon attacks the first and does nothing about the second. Each Jalapeño package reportedly carries six HBM4 stacks for 216 GiB and 15.4 TB/s of bandwidth, which means OpenAI is bidding for the same constrained supply as everyone else — and OpenAI's 10GW Broadcom commitment makes it one of the largest new sources of HBM4 demand, arguably tightening the very shortage it is exposed to. So the correct read is that Jalapeño improves OpenAI's position against one cost driver, worsens the industry's position on the other, and leaves the floor under inference prices roughly where it was.
Should this change which model vendor we standardise on?
Not by itself, and treating it as a vendor-selection input would be premature. The decision-relevant facts are still model quality, price, latency, data terms and switching cost, all of which are observable today, whereas Jalapeño's effect on any of them is a 2027 question resting on numbers one party supplied. There is a narrower and more defensible conclusion. Every serious lab now has a silicon hedge — Google with its own TPU line, Anthropic across AMD and TPU capacity, OpenAI with Jalapeño — which means the old concern that a single GPU supplier's pricing could squeeze every model vendor simultaneously has weakened considerably. That is a mild argument for expecting the price war to persist rather than resolve, and therefore for keeping switching costs low rather than signing long exclusive commitments. Keep a second frontier model green in evaluation and treat vendor choice as a routing decision you can revisit, which is the same conclusion the pricing evidence already supported.
What is the realistic risk to Nvidia here?
Slower and narrower than a single benchmark post suggests, but directionally real. The immediate threat is not that Nvidia loses OpenAI as a customer — OpenAI continues to buy Nvidia systems at very large scale and Jalapeño is explicitly an inference part, leaving training on merchant GPUs. The threat is to mix and therefore to margin: inference is the larger and faster-growing share of total AI compute over time, and it is precisely the workload most amenable to a fixed-function ASIC, because serving a transformer is a narrower problem than training arbitrary architectures. If the largest inference buyers each move a meaningful fraction of serving onto in-house silicon, Nvidia keeps the revenue growth but faces steady pressure on pricing power in its highest-volume segment. That is a multi-year erosion story rather than a cliff, and the key uncertainty is execution: Jalapeño's gen-1 is compared against Blackwell-generation systems while Nvidia's Vera Rubin generation is already shipping, so the relevant question is whether OpenAI's roadmap holds its advantage against a moving target.
Sources
- OpenAI — Jalapeño's first results show industry-leading speed and efficiency in AI inference (25 August 2026)
- SemiAnalysis — OpenAI Jalapeño: Better Than Nvidia Blackwell (InferenceX methodology and caveats)
- TechCrunch — OpenAI's Jalapeño chip is built for fast inference at scale, benchmarks show (25 August 2026)
- Tom's Hardware — OpenAI's 700W Jalapeño ASIC outpaces 1,400W Nvidia flagship GPU
- CNBC — OpenAI's Jalapeño AI chip brings new 'threat' to Nvidia margins as custom silicon gains ground (26 August 2026)
- TrendForce — OpenAI debuts Jalapeño AI inference chip, with Samsung reportedly supplying HBM4
- The Register — OpenAI's upcoming Jalapeño chip looks like it'll be an inference beast
- OpenAI and Broadcom unveil LLM-optimized inference chip (original Jalapeño announcement, 24 June 2026)
Related tool reviews
Questions or corrections? Email Pick Right. Want the full list? See all news.