OpenAI used its own LLMs to write Jalapeño chip benchmark code, lifting DeepSeek MLA kernel performance from 0.31% to 88.94% of ceiling in about 40 hours
- When the first Jalapeño chips came back from the foundry in May, OpenAI pointed internal AI models at writing benchmark software; on DeepSeek's multi-head latent attention kernel benchmark, performance climbed from 0.31 percent of the theoretical ceiling to 88.94 percent in roughly 40 hours, a result the team says is repeatable.
- OpenAI's Ho says all schedule assumptions now rest on this capability, so the gap between first silicon and production ramp can shrink.
- The article says Jalapeño cuts end-to-end latency from prompt to last token by up to 3.6x, without naming the baseline or the metric, which commenters call impressionistic math.
- Commenters call the title clickbait: nothing in the piece shows AI doing creative chip design, only software development inside a project run by people, including hires from Google's TPU team.
- OpenAI's LLM-assisted flow does not loosen the foundry bottleneck, so cheaper chips or cheaper RAM are not a near-term outcome.
Hacker News opinions
I grow jalapeños, so the branding really irks me. Jalapeño was already a Java VM at IBM, and now it is a chip too. Why not invent a new word instead of hijacking an existing one?
Guess how electrical engineers feel about the word transformers. Every field borrows names from somewhere, we are not special.
I am a licensed architect. Welcome to our hell of the last forty years of this.
Aw, I was expecting more details, but this is basically a rehash of what they unveiled a month ago.
Having worked with people doing bringup of specialized chips, I am awed. Going from 0.31 percent of the theoretical ceiling to 88.94 percent in about 40 hours is wild, and Ho says it repeats. Back in the day you wrote the code before the chip came back.
That 3.6x latency claim is impressionistic math. Is it 4.6 against 1.0, or 3.6 against 1.0? On spectrum.ieee.org I would expect a precise metric and baseline, and which latency is even being measured.
At some point someone will use an LLM to design an Apple M series competitor.
They won't, they'd need an ARM architecture license. And production CPU design is way more than RTL. Getting the power/performance/area numbers competitive takes heavy physical design optimization, and LLMs are not suitable for that work.
Everyone is still bottlenecked on foundries, not designs. That is why we are not getting cheap chips or cheap RAM any time soon.
Seems obvious OpenAI is just hyping their models so chip developers adopt them and they learn from that IP. Nothing in the article says AI did anything creative, it is software development inside the project.
Whatever happened with the Apple lawsuit? Apple does not have datacenter class accelerators, so there is nothing to steal in that space.
It is surprising that recursive self improvement looks more plausible than it did in 2023, but a 20 month turnaround for a chip is not breakneck speed. Physical manufacturing stays a hard obstacle.
Congrats to the former TPU team. I was surprised to see XLS in there until I remembered Chris went there a couple of years ago.
OpenAI should figure out how to build a lithography machine so ASML does not have a monopoly.
The Chinese have been working on EUV for a while already, so that race is not new.