OpenAI's 'Jalapeño' chip beats Nvidia's Blackwell on efficiency in first benchmarks — released one day before Nvidia reports earnings
OpenAI and Broadcom published test results showing their 3-nanometre inference processor delivering up to 1.9 times more throughput per kilowatt than Nvidia's flagship rack systems. Analysts are calling custom silicon the biggest competitive threat Nvidia has faced — and the timing, 24 hours before Nvidia's earnings call, looks anything but accidental.

OpenAI picked its moment. On Tuesday, one day before Nvidia was due to report quarterly earnings, the company published the first benchmark results for Jalapeño, the custom inference chip it has spent the past year designing with Broadcom — and the numbers landed exactly where Nvidia's shareholders did not want them. In testing overseen by the research firm SemiAnalysis, which visited OpenAI's labs, Jalapeño delivered 1.5 to 1.9 times more throughput per kilowatt than Nvidia's GB200 and GB300 rack systems, with end-to-end latency 1.7 to 3.6 times lower, beating the Blackwell generation on performance per watt in nearly every scenario tested.
Nvidia reports its results after Wednesday's closing bell. The juxtaposition is the story: the most valuable company in the world will spend Wednesday evening defending gross margins of roughly 75 per cent, hours after its single most famous customer demonstrated a chip designed to stop paying them.
A pepper aimed at a margin
Jalapeño is not a graphics processor and OpenAI is careful not to call it one — the company brands it an "Intelligence Processor," a chip built from the ground up for one job, running large language models in inference, the serving side of AI where every ChatGPT answer is generated. It was announced in June, and the design was co-developed with Broadcom, fabricated by TSMC on its 3-nanometre process, integrated into boards and racks by Celestica, and — according to reporting from Korean outlets carried by TrendForce — supplied with HBM4 memory by Samsung.
The engineering is genuinely unusual. OpenAI says the chip went from initial design to tape-out in nine months, a cycle the company claims is the fastest ever achieved in high-performance semiconductors; the industry norm is closer to two years. OpenAI used its own AI models to accelerate the design work, which makes Jalapeño a recursive kind of product — a chip partly designed by the software it exists to serve.
Die-shot analysis by Tom's Hardware suggests a single compute chiplet of roughly 840 square millimetres, close to the physical limit of what EUV lithography can print, ringed by six stacks of high-bandwidth memory. Per package, the reported specification is 216 gigabytes of HBM4, 15.4 terabytes a second of memory bandwidth, and a 700-watt rating — against 1,200 to 1,400 watts for the Nvidia systems it was benchmarked against. OpenAI says sustained draw in testing stayed at or below 550 watts.
"Jalapeño was designed from the ground up for LLM inference," said Richard Ho, who leads OpenAI's hardware programme. "Based on early testing, Jalapeño will efficiently execute our most important workloads close to the hardware's theoretical limits."
The caveats Nvidia will cling to
There are real asterisks, and SemiAnalysis — whose own testing produced the headline numbers — was the first to attach them. Jalapeño exists today as engineering samples running in a lab. Nvidia's next-generation Vera Rubin systems, which also use HBM4, are shipping to paying customers now. "Jalapeño is really competing against chips like Rubin," the firm wrote, calling its own Blackwell comparison "somewhat incomplete and unfair." The efficiency figures are also normalised to published power ratings rather than measured wall draw, a choice that can flatter either side depending on how the chips behave under load.
The chip is an inference specialist, not a training chip. For the frontier work of building new models, TrendForce analyst Fion Chiu told CNBC, "we believe Nvidia GPUs will remain important given their broad programmability, performance, software ecosystem, and ability to handle a wide range of workloads."
And scale is not a press release. OpenAI says Jalapeño will be deployed in its infrastructure by the end of the year, with Broadcom chief executive Hock Tan promising "gigawatt scale data centers with Microsoft and other partners beginning in 2026," a second generation possibly taping out within months and a third already in development. Until those racks are serving traffic, Nvidia still sells every accelerator that matters.
Why analysts used the word 'threat'
The reason Tuesday's benchmarks rattled analysts is not that one chip beat another in a lab. It is what the result proves is possible. "A hyperscaler-designed chip can now match or beat Nvidia's Blackwell-class GPUs on inference efficiency," Adrien Sanchez, a technology analyst at Yole Group, told CNBC — a direct "threat to Nvidia's inference margins, which is the field growing the most at the moment."
This is the biggest competitive threat to Nvidia, as about half the capital expenditure on AI infrastructure comes from hyperscale cloud providers who either have a custom chip program or could reasonably have one. — Alexander Harrowell, senior principal analyst, Omdia
Harrowell's firm now expects custom chips like Jalapeño to exceed GPUs in unit volume by 2028, though revenue parity will take far longer because GPUs cost considerably more. In a gigawatt-scale deployment, he added, Jalapeño's efficiency "would save power, cooling, and power distribution infrastructure, and contribute a lot to their unit economics" — which is the entire game when OpenAI has told investors it expects to spend some $600 billion on infrastructure by 2030.
The rest of the industry is already there. Google's seventh-generation TPU, Ironwood, is being deployed by Anthropic at a scale the industry had never seen for custom silicon — the company has disclosed more than one million of the chips serving its Claude models. Microsoft announced its Maia 200 in January. Meta agreed in April to deploy a gigawatt of Broadcom-based custom accelerators. Industry trackers estimate hyperscalers will field roughly 1.9 million custom accelerators in 2026. Broadcom, which sits behind several of those programmes, has told investors it expects more than $100 billion in AI revenue in its 2027 fiscal year — and marked its confidence in August by granting Google warrants tied to as much as $120 billion of future custom-chip purchases through Marvell's ecosystem.
There is a second-order consequence hiding in the specification sheet. Six stacks of HBM4 per package, multiplied across gigawatt-scale deployments, makes Jalapeño a major new claimant on the world's most constrained commodity: high-bandwidth memory. Samsung, SK Hynix and Micron have capacity essentially committed through 2027, and Tom's Hardware reports Nvidia has already tested pared-down configurations of its Rubin Ultra parts with as little as 192 gigabytes of memory. Every wafer of HBM4 that goes into an OpenAI inference rack is a wafer that does not go into an Nvidia one — the competition for supply starts before the competition for customers does.
The strangest customer relationship in business
What makes the OpenAI–Nvidia story unlike any normal supplier dispute is that the two companies are financially entangled to a degree that has no precedent. On August 17, an SEC filing revealed that Nvidia will provide up to $105 billion in credit support for OpenAI's enormous datacenter campus in Pike County, Ohio — a guarantee that triggers only if OpenAI defaults, and a site that will run exclusively on Nvidia compute. Nvidia is separately finalising an equity investment of roughly $30 billion in OpenAI, the successor to a $100 billion letter of intent that stalled last winter and that chief executive Jensen Huang later said "was never a commitment."
So Nvidia is guaranteeing the debts of the customer that is building the chip designed to displace it — while that customer hedges in every direction at once. OpenAI has a 10-gigawatt custom-accelerator agreement with Broadcom signed last October, of which Jalapeño is the first fruit; a two-gigawatt commitment to Amazon's Trainium; a deal with the startup Cerebras that OpenAI has confirmed exceeds $10 billion; and its Oracle rental contract, worth some $300 billion over five years, remains the largest single compute deal ever signed.
"The world is moving to a compute-powered economy," OpenAI president Greg Brockman said in the announcement. "By designing more of the stack ourselves, we can serve more intelligence with greater efficiency and keep pushing advanced AI toward broader access."
What Wednesday night decides
The market's first verdict was a shrug — Nvidia stock closed Tuesday at $213.83, up 0.37 per cent, suggesting investors either doubt the benchmarks or believe the AI pie is growing faster than any one chip can carve it. The second verdict comes Wednesday evening, when Nvidia guides on a quarter analysts expect to land around $91 billion in revenue at roughly 75 per cent gross margin, and Huang inevitably fields the question CNBC's Investing Club says is coming: what happens to that margin when your biggest customers become your competitors?
Hock Tan gets his turn on Broadcom's call on September 2. Between the two of them, the market will learn whether Jalapeño is a lab curiosity with excellent public relations — or the first proof that the most profitable product line in the history of the semiconductor industry has a ceiling, and that the customer paying for most of it just found the door.
Sources: OpenAI and Broadcom announcements, SemiAnalysis published benchmarks, CNBC, Tom's Hardware, TrendForce, Bloomberg, Fortune and SEC filings. Figures as reported August 26, 2026.
