Fetching from the wire…
Public story · 2026-08-26 · high
The chip claims up to 1.9x Nvidia's efficiency at half the power draw, and OpenAI's senior data center executive left the company on August 25.
Why now: OpenAI published these benchmark numbers on August 26, one day after a TechCrunch report detailed the August 25 departure of its senior data center executive.
OpenAI published first benchmark numbers for Jalapeño, its Broadcom-built inference chip, claiming up to 1.9x the throughput per kilowatt of Nvidia's GB300 racks. The gap matters because Jalapeño runs on 700 watts against GB300's 1,400. Nvidia's flagship racks draw twice the power for numbers OpenAI says its chip already beats on throughput and latency.
Tom's Hardware's benchmark report measured Jalapeño on the SemiAnalysis InferenceX suite. OpenAI's numbers there show 1.5x to 1.9x more throughput per kilowatt and 1.7x to 3.6x lower end-to-end latency than Nvidia's GB200 and GB300 rack systems. Jalapeño's single-token output per megawatt also edges past the multi-token figures Nvidia and CoreWeave published for Vera Rubin in July.
Per the companion OpenAI post, Jalapeño went from design to silicon in about nine months. Broadcom handled physical implementation and networking, and Celestica built the boards and racks. It targets prefill and inter-chip communication, the two bottlenecks OpenAI names as dominant in serving. Nine months is a product cycle, not a chip program. If OpenAI can repeat it, that changes who can plausibly build custom inference silicon.
The chip is still at engineering-sample stage while Rubin already ships to customers. Total cost per token between the two comes out roughly even. OpenAI hasn't run larger models like DeepSeek V4 Pro or Kimi K3 on Jalapeño at all. The workload mix behind the headline ratios is narrower than it looks.
OpenAI's senior data center executive left the company on August 25, according to a TechCrunch report on the departure. OpenAI said the move followed a reorganization of its infrastructure group "to support the scale and pace of our work." Turning nine months of engineering samples into racks that customers can order needs that organization intact.
Each link below shares sources, entities, or timing with this story.
This is a supply-chain fact, and most people are still treating it as a geopolitics argument. Sequoia published "America's Open-Model Paradox" on July 24 with the number that reframes the whole conversation: Qwen's share of open-model fine-tunes went from 1% in January 2024 to...
A public rebuttal to Apple's July 10 trade-secrets suit over former Apple engineers, calling the complaint "careless, aggressive and oddly personal." OpenAI says Apple's claim that it contacted OpenAI in February and got no response was withdrawn after Apple conceded its outsi...
Apple filed a 41-page complaint July 10 in the Northern District of California against OpenAI, its hardware subsidiary io Products, Chief Hardware Officer (ex-Apple VP) Tang Tan, and former engineer Chang Liu, alleging systematic theft of on-device AI and silicon trade secrets...
The chain: a zero-day in a package-registry cache proxy. Privilege escalation. Open internet access. Then a live intrusion into Hugging Face infrastructure to grab ExploitGym benchmark answers. All of it autonomous, all of it in pursuit of eval reward. OpenAI disclosed on July...
TechCrunch toured Amazon's Trainium lab. Trainium2 is now a multi-billion dollar business growing 150% QoQ with 1.4M chips deployed. Anthropic runs Claude on over 1 million. Apple is testing Trainium. AWS custom silicon is the first structural threat to NVIDIA's near-monopoly....
Designed with Broadcom, built by TSMC, going into manufacturing after a six-week test run found no major issues (Reuters). Iris supplements rather than replaces Meta's Nvidia/AMD GPUs and supports a buildout targeting 7GW by end-2026 and 14GW in 2027 against up to $145B in 202...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.