OpenAI's Jalapeño Chip Smokes Nvidia's Blackwell in Inference
OpenAI Jalapeño: Better than Nvidia Blackwell

At Hot Chips, OpenAI unveiled Jalapeño, an inference chip built with Broadcom that outperforms every Nvidia, AMD, and Google chip in SemiAnalysis's InferenceX benchmarks. Despite a rapid 16-month development cycle, the chip delivers industry-leading performance per watt, even beating Nvidia's upcoming Rubin on token throughput per megawatt. SemiAnalysis details the architecture, software, and benchmark results, noting that Jalapeño is a general-purpose inference chip, not specialized for OpenAI models, and that its performance is achieved without speculative decoding.
Jalapeño smokes every other chip.
- Traster
This semi-analysis article reads a lot more like an OpenAI press release than a real analysis. And to be honest some of the statements seem like just straight up lies - they initially claim they were invited to benchmark it, and then half way down switch to claiming that OpenAI provided all the numbers. This really kind of sucks, because I want to read actual detailed nuanced and credible analysis of what's happening in the industry and it doesn't seem like you can trust this as far as you can throw it.
- mchusma
I think they talked about this being general purpose chip but I would think that Anthropic/OpenAI are at the scale now they could bake LLM weights into chips themselves.
For example, GPT Sol baked into a custom chip run for $100M that runs 10x as fast and 10x as cheap should pay for itself as long as the chip is useful for long enough.
While 2 years ago nothing was useful more than 1 year long, there are many older models in use now (e.g. Haiku 4.5, GPT-OSS 120b), and I expect this trend to continue.
I know this is what Taalas was doing (acquired by AMD), here was their demo, https://chatjimmy.ai/ which is based on Llama 3.1 8B. It feels like this should start to happen soon.
- iFire
So out of all the inference only chips which ones can I buy?
The only report is a smartnic fpga from Alibaba where we take an onnx design and write our own.
https://essenceia.github.io/projects/alibaba_cloud_fpga/
On my M2 Pro Mac Mini the ANE only allows 2 gigabytes compared to the Metal GPU which can use the system ram.
Currently playing with https://www.asus.com/motherboards-components/ai-accelerator/... which is a 4bit, 8bit and 16 bit ai inference chip with 8 gigabytes of ram.
The UGen300 has the Hailo-10H chipset.
The ASUS Store price for the ugen300-usb-8g costs $365.00 Canadian dollars.
- g00afthrowaway
Funny semi analysis has credibility here of all places. The founder is well-known in the hardware circle to be a black-market information trader.
It works like this:
1. Founder befriends undergrad interns/graduate student interns, buys them gifts, invite them to dinner/yacht/house/vc parties etc, or pays them to write articles
2. Founder extracts insider information out of these interns
3. Founder sells this information to companies paying "consulting" fees
- epistasis
It's so funny to see FP4.... I remember 20 years ago being asked what sort of HPC we needed in genomics, and the answer was basically, "lower precision, faster" for the stuff I was working on. But FP4 is, well, almost comical.
One thing not on that comparison table: die size. If I'm understanding that correctly, it's about the same as the Rubin, but at 1/3 the number of NVFP4 PFLOPs. (The text disagrees with the table, I'm taking the table as truth, perhaps that's wrong...)
- corford
These nascent inference chip efforts are reminding me of the early 3dfx / riva / mach / powervr days. Will be interesting to see if inference chips are here to stay and, if so, who the eventual dominant player(s) will be
- fraboniface
I hadn't seen the token/Joules comparison with human speech before. Humans are still 22x more efficient, which is not that far considering the rate of progress in this area.
- jimmySixDOF
I love how now you have to consider the possible s** posting motivation behind analysis of a trillion dollar industry being conducted at a world-class level by a bunch of ex Reddit and 4Chan adjacent mods -- it's one of the best stories in AI that SemiAnalysis is not cut from the same cloth as Gartner McKinsey et al
- anthonypasq
Continued hardware improvements really make it hard for me to believe token prices will not continue to plummet.
- blt
It can't come soon enough that AI workloads get their own specialized hardware to free up the general-purpose devices for general-purpose usage again.