Find Our Latest Video Reviews on YouTube!
If you want to stay on top of all of our video reviews of the latest tech, be sure to check out and subscribe to the Gear Live YouTube channel, hosted by Andru Edwards! It’s free!
Thursday August 27, 2026 9:29 pm
OpenAI Built Its Own AI Chip, Jalapeño, and It’s Beating Nvidia on Power
Posted by Andru Edwards Categories: Corporate News, Artificial Intelligence

You know that pause when you ask a chatbot something and the answer sits there, one word at a time, while you stare at it? That pause is inference, and it's the most expensive thing in AI right now. OpenAI just showed off a chip built to kill it.
The chip is called Jalapeño. OpenAI published its first benchmark numbers at the Hot Chips conference this week. Against Nvidia's Blackwell systems, OpenAI says Jalapeño does 1.5x to 1.9x more work per watt and responds 1.7x to 3.6x faster end to end.
What the numbers mean
The tests came from InferenceX, a public benchmark from the research firm SemiAnalysis, run across three open models: GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. SemiAnalysis says it verified some of the runs in OpenAI's lab. For the most interactive workloads, the kind where a person sits there waiting on a response, OpenAI claims 2.1x to 4.1x better performance.
Jalapeño is a 700W part going up against Nvidia systems rated at 1,200W and 1,400W. AI data centers are increasingly limited by how much electricity they can physically get, so work-per-watt decides how much AI a company can serve.
A lot of coverage has said Jalapeño is tuned specifically for OpenAI's own models. It isn't. It's a general purpose inference accelerator that runs whatever you throw at it, and OpenAI showed SemiAnalysis the chip running Doom, ported over using Codex prompts. It handles inference only, so it doesn't train models.
The caveats
SemiAnalysis called the Blackwell comparison somewhat incomplete and unfair. Jalapeño uses newer HBM4 memory, so the honest matchup is against Nvidia's newer Vera Rubin platform. Jalapeño still squeezes out more output tokens per megawatt there, but on total cost of ownership per token, the two come out roughly even.
Jalapeño also posted its results without multi-token prediction enabled while the competing chips ran with it, which cuts in OpenAI's favor. And it has no CUDA, which has kept Nvidia's moat filled for a decade. OpenAI has been explicit that it will keep buying Nvidia for both training and inference. Deployment is small scale by the end of this year, with the bigger rollout in 2027.
What it means for you
Inference cost is what forces AI companies to ration the good models behind subscription tiers and usage caps. Cheaper inference is what eventually gets the fast, smart model to the free tier.
OpenAI also says it used its own models to design this chip, compressing design-to-tapeout to around nine months and, for some blocks, generating optimized software kernels that outperformed human-written ones by 1.5x to 1.8x. AI is now helping design the hardware that runs AI.
Nvidia's data center business is still growing at a pace one customer's in-house chip cannot dent. But a first-generation part beating Blackwell on efficiency puts a number on how replaceable Nvidia is for inference work, and that number is no longer zero.