Close-UpTuesday, August 2510 Min Read
OpenAI's Chip Beat Nvidia's: The Edge Is in the Watts
OpenAI published the first measurements of Jalapeño, the chip it built with Broadcom, claiming up to 1.9x the tokens per kilowatt of Nvidia's GB300. Most of that advantage comes not from the chip but from a power draw that is half as large.
On Tuesday, August 25, OpenAI published measured results for the first processor it designed itself. Against Nvidia's GB300, the chip called Jalapeño came out ahead on two counts: tokens produced per kilowatt consumed, and the time it takes a response to come back. The published ranges are 1.5x to 1.9x and 1.7x to 3.6x respectively.
The same day, Nvidia closed up 2.19% at $213.05.
Put those two sentences side by side and the first explanation that comes to mind is that the market missed the news. Read the measurement closely and a second explanation appears: the market saw it, and separated the numerator of the ratio from the denominator.
By the Numbers
700W
Jalapeño's rated power
1,400W
Rated power of the GB300 it was measured against
1.9x
Highest claimed advantage per kilowatt
10 GW
Capacity to be built with Broadcom through 2029
A Processor Designed in Nine Months
Jalapeño itself is not new. OpenAI and Broadcom introduced the chip on June 24, and no performance figures were given that day. What was shared was physics: the compute die measures roughly 840 mm², close to the 858 mm² reticle limit of EUV lithography, with six memory stacks around it, fabricated on TSMC's 3nm-class process. The companies said the design went from concept to tape-out in nine months, which they called the fastest ASIC development cycle achieved in high-performance semiconductors. An ASIC is a chip built for one job rather than general-purpose work, and Jalapeño's one job is inference on large language models — running a trained model. It does not train them.
Broadcom supplied the silicon implementation and Tomahawk networking; Celestica handled board and rack design. The frame around all of it is the agreement announced in October 2025 to deploy 10 gigawatts of OpenAI-designed accelerators by 2029. Broadcom chairman Hock Tan said at the launch:
"This is just the beginning of a multi-generation roadmap."
Whose Hands Jalapeño Passes Through
- 01DesignOpenAI
- 02Silicon and networkingBroadcom · Tomahawk
- 03FabricationTSMC, 3nm class
- 04MemorySix HBM4 stacks
- 05Board and rackCelestica
The only thing missing in June was the numbers. August supplied them.
Who Ran the Measurement, and Where
The figures come from InferenceX, SemiAnalysis's benchmark suite. The tests ran on three open models: GPT-OSS 120B, DeepSeek R1 670B, and Moonshot AI's trillion-parameter Kimi K2.5. The workload was an 8,000-token prefill with a 1,024-token decode.
Where the measurement happened matters as much as the result. SemiAnalysis writes that the runs took place in OpenAI's lab with OpenAI engineers present, that the results were provided to them by OpenAI, and that while they watched the runs in person they did not execute the full suite themselves:
"we verified the InferenceX runs in person in the lab"
None of this makes the result wrong. It makes it a different kind of evidence than a measurement reproduced in an independent lab, and disclosing that difference is the job of whoever publishes the benchmark. Here it was disclosed.
Two more technical details. First, the 8,000/1,024 shape is, in SemiAnalysis's own words, a "much easier workload to tune for" than their long-context, multi-turn suite. Second, Jalapeño was measured using single-token prediction, while the competing results had multi-token prediction enabled. The same analysis notes that speculative decoding delivers a 3x-plus improvement in cost per token. The asymmetry, in other words, also runs against Jalapeño.
Where the Advantage Comes From: The Denominator
This is the mechanism. "Performance per kilowatt" is a division: tokens produced over power consumed. There are two ways to raise a ratio — grow the numerator, or shrink the denominator. Jalapeño is rated at 700W; the GB300 at 1,400W. The denominator is exactly half.
This is not an accusation. In a data center the scarce input is not chips but electricity: grid interconnection is a queue measured in years, and cooling capacity per rack is a fixed ceiling. In that world, performance per kilowatt is the correct denominator. But a reader who sees "beats Nvidia" in a headline understands that one chip outperforms another chip. The measurement does not say that.
The choice of denominator moves the answer again. The results were normalized to each side's published TDP, and Jalapeño's measured sustained draw in the lab was 550W, not 700W. The same analysis notes that computing on all-in facility power narrows the gaps. In a benchmark, which number goes under the line is often more decisive than the number above it.
The Range Behind the Claim
The figures carried in headlines were taken from the top of the range. The bottom is part of the same measurement.
The Wrong Opponent
The GB300 that Jalapeño was set against is a 2025 Nvidia design. Its memory is HBM3E: 288GB of capacity at 8 TB/s of bandwidth. Jalapeño uses HBM4 — 216 GiB of capacity at 15.4 TB/s. Less capacity, nearly twice the bandwidth. In the token-generating phase of inference the bottleneck is bandwidth, not capacity, which makes this a comparison between two memory generations.
Nvidia's HBM4 part is Vera Rubin. SemiAnalysis calls the comparison against Blackwell "somewhat incomplete and unfair" in its own text and says Vera Rubin is the proper baseline. In measurements it published in July 2026, Vera Rubin NVL72 delivered roughly 5.4x the throughput per megawatt of the GB300 NVL72 at the highest interactivity setting and around 4x at a middle setting. The same analysis finds that once divided by total cost of ownership, Vera Rubin and Jalapeño come out head-to-head.
| Jalapeño | GB300 | Vera Rubin | |
|---|---|---|---|
| Memory generation | HBM4 | HBM3E | HBM4 |
| Memory capacity | 216 GiB | 288GB | not disclosed |
| Bandwidth | 15.4 TB/s | 8 TB/s | 2.8x Blackwell Ultra |
| Rated power | 700W | 1,400W | not disclosed |
| Workloads | inference only | training and inference | training and inference |
The last row of that table goes missing from most commentary. Jalapeño cannot train a model. OpenAI's hardware lead said so plainly:
"Nvidia is a really good partner, and we continue to need a lot of Nvidia."
The Fifteen Months Between a Benchmark and a Shipment
For a benchmark result to register anywhere in the supply chain, the chip has to be built. According to SemiAnalysis, Jalapeño production ramps gradually through 2027, with most of the output currently scheduled for the fourth quarter of that year. The measurement was published in August 2026. Volume shipment sits fifteen months out on the calendar.
Another constraint is waiting in those fifteen months. According to press reports, the entire 2027 memory capacity of Samsung, SK Hynix and Micron has already been contracted. Every copy of Jalapeño wants six HBM4 stacks.
A rough upper bound is worth doing. If the full 10 gigawatts were accelerator power, it would take 14.3 million chips at 700W each, which comes to 85.7 million HBM4 stacks. The real number is meaningfully lower, because a large share of a data center's electricity goes to networking, host CPUs and cooling. Even so, the order of magnitude shows what one customer's custom chip means to a memory maker. The performance argument is being fought at the accelerator layer. The binding constraint sits one layer below it, in memory allocation.
Timeline
- October 2025OpenAI and Broadcom announce a 10-gigawatt collaboration through 2029.
- June 24, 2026Jalapeño is introduced; no performance figures are shared.
- July 2026SemiAnalysis publishes its Vera Rubin NVL72 measurements.
- August 25, 2026The first Jalapeño benchmark results are published. Nvidia rises 2.19%.
- August 26, 2026Nvidia reports earnings after the close.
- Q4 2027Most of Jalapeño's production is scheduled for this period.
What the Market Actually Said
Nvidia rose 2.19% to $213.05 on August 25. Pinning that move on the Jalapeño story alone would be wrong; other things were moving the price that day. The Monday before, August 24, the Nasdaq fell 0.8% as semiconductor names sold off ahead of Nvidia's report. On Tuesday several brokerages raised their Nvidia price targets. The Tuesday gain is at least as consistent with a rebound inside an earnings week.
What can be said with confidence is narrower and more useful: the market did not price the published benchmark as an event that changes Nvidia's income statement. Sometimes the most informative data is not the move but its absence. When a piece of news is not priced, the market does not yet regard it as an order, a contract, or an allocation of capacity.
Nvidia reports after the close on August 26. The lines to watch are gross margin and the composition of data center revenue — competitive pressure from custom silicon, if it exists, shows up first in pricing.
The chart above is a picture of six months, not of one news item; it cannot tell you which move belongs to which headline.
This piece draws on the announcements published by OpenAI and Broadcom, on SemiAnalysis's analysis built on its InferenceX benchmark suite, on Bloomberg's August 25 report, and on Tom's Hardware's technical coverage. The performance figures were supplied by OpenAI, and the party publishing the benchmark stated that it did not independently run the full suite; that distinction is preserved throughout. Power and memory capacity figures for Vera Rubin have not been disclosed publicly. Nvidia's August 25 close was $213.05. This is not investment advice.