OpenAI's Jalapeño chip outperforms Nvidia Blackwell on inference throughput per watt

OpenAI has built its own AI inference chip, codenamed Jalapeño, and according to a detailed technical analysis by SemiAnalysis published on August 25, it delivers better throughput per megawatt than Nvidia's Blackwell accelerators on key large language model workloads. The disclosure marks the clearest signal yet that OpenAI intends to escape its dependence on Nvidia's supply chain.
What Jalapeño is and what it can do
Designed in-house and manufactured on TSMC's N3P process, Jalapeño is a weight-stationary systolic array chip with 700W TDP, 15.4TB/s of HBM4 memory bandwidth (at 10Gbps pin speeds), and 13.4 PFLOPs of MXFP4 compute. Each rack holds 128 chips; the system scales to 2,048 chips across 16 racks connected via Tomahawk 6 switches and optical circuit switches.
SemiAnalysis benchmarked the A0-stepping engineering samples against several major models. On DeepSeek R1, Jalapeño exceeded 700 tokens/second at single-user concurrency. On Kimi K2.5, it approached 700 tokens/second per user — representing a claimed 9x advantage over comparable Blackwell configurations. On GPT-OSS, it delivered 1,400 tokens/second per user. All figures were achieved without speculative decoding or prefill-decode disaggregation.
The CUDA moat may be cracking
The more consequential finding in the SemiAnalysis report may not be the raw performance numbers but the software story. The analysis notes that "OpenAI's software bring-up has progressed more quickly than Nvidia's," and states outright that "the CUDA moat is potentially dead" given the hardware-software codesign OpenAI is executing at scale. Kernel performance improved 2x in under two weeks during testing, and the team scaled from single-system configurations to full rack-scale operations in eight days.
Design began in mid-2024, with tape-out completed in November 2025 — a 16-month cycle that is fast for custom silicon at this scale. Production volumes are currently minimal, with most output scheduled for the end of 2027. The B0 stepping, which follows the tested A0, promises a 25% efficiency improvement.
Context and caveats
The comparison to Blackwell is, as SemiAnalysis itself acknowledges, "somewhat incomplete and unfair" — Jalapeño's HBM4 more directly competes with Nvidia's upcoming Vera Rubin platform than with Blackwell, which uses HBM3e. Testing covered only 8k-context, 1k-output benchmarks and did not include the AgentX suite that reflects production-scale agentic workloads. Larger models such as DeepSeek V4 Pro were not tested, and all performance numbers were provided by OpenAI and only partially verified independently.
Still, even with those caveats, the emergence of a competitive custom chip from OpenAI shifts the strategic calculus for the entire AI hardware market. Nvidia's data-center revenue has grown to more than $85 billion per quarter partly because large AI labs had no credible alternative. Jalapeño is the clearest sign that this is changing — even if broad deployment is still 18 months away.
Originally reported by SemiAnalysis. Read the original article for additional details.
View original source