OpenAI says its Jalapeño chip beats Nvidia's GB300 per watt
OpenAI's first published Jalapeño benchmarks claim 1.5 to 1.9 times more work per kilowatt than Nvidia's GB300, at half the power. The numbers are OpenAI's own.
4 min read

By the numbers
- more throughput per kilowatt, per OpenAI
- 1.5-1.9x
- lower end-to-end latency at the top of the range
- 3.6x
- HBM4 memory on each chip
- 216 GB
- from first design to tape-out
- 9 months
- OpenAI Jalapeño
- 700W
- Nvidia GB300
- 1400W
OpenAI has published the first benchmarks for Jalapeño, the inference chip it built with Broadcom, and says it beats Nvidia's flagship. The Register and TechCrunch reported the figures on 25 August 2026. OpenAI claims 1.5 to 1.9 times more AI work per kilowatt than competing systems, and 1.7 to 3.6 times lower latency. The chip is rated at half the power of Nvidia's GB300. If the numbers hold up outside OpenAI's own tests, the cost of serving a model just got a second supplier.
Inference is the work of running a trained model to answer requests, as opposed to training it. It is where most of the electricity in an AI product goes once the product has users. A chip that does more inference per watt cuts the biggest line on the bill.
The numbers OpenAI published
Tom's Hardware reports Jalapeño is rated at 700 watts against 1,400 watts for Nvidia's GB300. On that basis OpenAI claims up to 1.9 times the throughput per kilowatt. The Register adds that for ultra-low-latency work the advantage rises to 2.1 to 4.1 times, and estimates the chip's power use at 40% to 60% of competing GPU systems.
The single-chip specifications, per The Register:
| Measure | Jalapeño |
|---|---|
| Compute at MXFP4 | 13.4 petaFLOPS |
| Memory | 216 GB HBM4 |
| Memory bandwidth | 15.4 TB/s |
| Rated power | 700 W |
MXFP4 is a 4-bit number format, which packs weights tightly so more of the model fits in fast memory. HBM4 is high-bandwidth memory, the stacked memory that sits next to the chip. A full rack of 128 chips reaches 1.7 exaFLOPS of 4-bit compute, 27.5 TB of memory, and nearly 2 petabytes per second of bandwidth.
How it gets the speed
The design bet is on moving data less. Richard Ho, OpenAI's head of hardware, told TechCrunch: "We designed Jalapeño to minimize data movement and communication delays." The Register reports a large on-chip SRAM cache and careful placement of the KV cache, the store of past tokens a model consults on every step.
That matters because inference has two phases. Prefill reads the whole prompt at once, and decode produces the answer one token at a time. Decode is the slow part, and it is bound by how fast the chip can fetch what it already knows. Keeping that data close is what the latency figures reflect.
The chip was also built unusually fast. The Register reports 9 months from first design to tape-out, the point where a design is frozen for manufacturing, and describes it as "designed (in part) by AI, for AI". TechCrunch confirms OpenAI's own models assisted in the design.
When you could actually use it
Not soon. TechCrunch reports OpenAI expects to deploy Jalapeño in "very small volumes" at the end of 2026, with wider deployment in 2027. The project was first announced in October 2025, and TechCrunch describes it as a multigenerational platform, so this is the first chip of a line.
Ho was direct about what he thinks the results mean. "The bottom line is that the results show a very, very significant performance advance over state of the art," he told TechCrunch. He added that Jalapeño "can serve more AI work per unit of power, while also returning responses more quickly".
What this means for developers
Read every figure with one fact in front of you: the numbers come from OpenAI. No independent lab has reproduced them yet, and the comparison is against Nvidia's published power rating. A vendor's first benchmark of its own chip is a claim, not a measurement. Treat it as a reason to watch, not a reason to re-plan.
If you run inference at scale, the metric to steal is work per kilowatt. Most teams still budget by GPU count. Power is the constraint that actually caps a data centre, and it is the axis OpenAI chose to fight on. Start tracking tokens per kilowatt-hour for your own workload now, so you can evaluate any chip, including this one, when it ships.
You will not buy a Jalapeño. The reports describe a chip built for OpenAI's own fleet, in small volumes at first. The way it reaches you is through OpenAI's API pricing and latency, in 2027 at the earliest. If your product depends on fast decode, watch OpenAI's per-token latency numbers next year for the first real signal.
The nine-month tape-out is the part worth remembering. A frontier lab designed a competitive inference chip in under a year, with its own models helping. If that holds, custom silicon stops being a five-year bet only the biggest companies can make. The defensible position for Nvidia becomes software and supply, not raw efficiency.
Sources
Related articles

Nvidia's Groq 3 LPX inference chip enters full production
Nvidia put its dedicated inference accelerator into full production, splitting agent workloads so GPUs handle context and the new chips handle token generation.

Fujitsu MONAKA ships 144 Armv9 cores in November
Fujitsu's 144-core MONAKA goes on sale in November, with 2nm cores stacked over 5nm cache and GCC 15 support already upstream.

OpenAI Agents API adds computer use in a hosted browser
OpenAI's Agents API, in public beta since September 10, gained computer use at DevDay on September 29, 2026: agents can drive an OpenAI-hosted browser.
The daily brief
Three to five stories a day, and what each one means for the people who build software. Free, no spam.