SemiAnalysis says OpenAI's Jalapeño inference ASIC beats Blackwell on throughput per MW
- SemiAnalysis says its lab-run InferenceX tests put OpenAI's Jalapeño ahead of Nvidia Blackwell in token throughput per all-in utility MW across nearly all tested scenarios, despite Jalapeño running without Multi Token Prediction while comparison configurations used it.
- OpenAI and Broadcom began Jalapeño design in mid-2024 and reached manufacturing tape-out in about 16 months; the chip program was publicly unveiled in June and the chip was announced at Hot Chips.
- The article describes Jalapeño as a general-purpose LLM inference ASIC rather than an OpenAI-model-specific accelerator, saying it ran multiple open models and workloads in the InferenceX suite with OpenAI engineers present.
- Jalapeño uses HBM4 and is positioned against flagship Nvidia and AMD GPUs; SemiAnalysis says its advantage spans both low-latency and high-throughput inference rather than one operating point.
Hacker News opinions
I find the speed of the chip program most interesting. If LLMs are helping engineers design accelerators, I wonder whether general-purpose chips will start improving much faster too.
Hardware gains like this make it hard for me to see token prices doing anything but fall.
Cheaper tokens do not help much if models simply burn more tokens on the same tasks.
Lower token prices need environmental and local regulation. Data centers still bring generator pollution, noise, water use, power-line projects, and pressure on nearby residents.
There is pressure on token prices from every direction. I would need to see collusion or regulatory manipulation to believe costs will not fall.
This is a proprietary accelerator built by the company selling tokens. Why assume OpenAI will pass efficiency gains to customers rather than keep the margin?
I am not convinced it is better than Rubin. The comparison makes Jalapeño look competitive, not clearly ahead.
Jevons paradox may apply. More efficient hardware cuts near-term token costs, but lower prices can create far more demand and raise total consumption.
Designing an ASIC is far easier than making leading-edge memory. DRAM has heavy patent barriers, huge R&D and fabrication costs, difficult process development, and high failure rates.
The FP4 numbers are funny after years of treating half precision as an aggressive tradeoff. The table appears to put Jalapeño near Rubin in die size but at about one-third of the NVFP4 PFLOPs, and its text may disagree with that table.
I like SemiAnalysis, but I would be cautious about its benchmarks. Critics say its scripts have counted numerators and denominators incorrectly and call the operation slipshod despite its expensive subscriptions.
I do not think inference becomes equally accessible just because it commoditizes. This looks more like oil, where a few companies have the scale to produce at the lowest cost and big labs can approach the energy cost of serving models.
Models are not interchangeable commodities like steel. They have distinct strengths and weaknesses, which is why users still argue over writing quality and model behavior.
Nvidia currently takes several times a chip's manufacturing cost, in my view. Competition and inference commoditization should cut that hardware premium and may eventually affect training costs too.
The human-speech token-per-joule comparison is incomplete. A brain's roughly 20 W supports everything else a person does, while these chips only emit tokens, and the output quality is not equivalent.