• STICKY POST

Find Our Latest Video Reviews on YouTube!

If you want to stay on top of all of our video reviews of the latest tech, be sure to check out and subscribe to the Gear Live YouTube channel, hosted by Andru Edwards! It’s free!

Tuesday August 25, 2026 12:47 pm

Nvidia’s Groq Chip Is in Full Production, and It Exists to Fix One Annoying Thing


nvidia groq 3 lpx

When you hand an AI coding agent a task and it sits there for eleven seconds before doing anything, you're feeling a very specific bottleneck. Not training. Not raw compute. Token generation speed, which is how fast a model produces its answer one piece at a time. Nvidia just put a chip into full production that does almost nothing else.

Announced at Hot Chips 2026, the Groq 3 LPX is an inference accelerator that extends Nvidia's Vera Rubin platform. It's also the first real product to emerge from Nvidia's $20 billion Groq acquisition, and Nebius is the first AI cloud signed up to run it.


What it's actually for

Nvidia splits the work in two. Rubin GPUs handle context processing, the heavy job of reading everything you fed the model. The LPX units handle decoding, the latency-sensitive phase where tokens come out one after another. Racks scale to as many as 256 LPX accelerators.

The number Nvidia is leading with: 3,400 output tokens per second running Gemma 4 31B with a 100,000-token context, in Artificial Analysis benchmarking. Nvidia says that's four times the nearest alternative platform and the fastest result ever recorded for that model.

Why long context is the whole point

Agentic systems burn tokens differently than chatbots do. An agent reads files, writes code, runs it, reads the error, and tries again. Each of those is a full generation pass, and hundreds of them stack up inside a single task. Slow generation doesn't just make one answer late. It compounds across the entire loop until the agent stops feeling worth using.

That's the case Nvidia is making, and it's a reasonable one. It also happens to be very good for Nvidia's business. Clouds can charge premium rates for low-latency tokens, which is precisely the argument that gets a chip like this bought.

Who's buying

Nebius is first, bringing LPX into its Token Factory inference platform, with the pitch that developers get the speed through the same API they already use and skip any migration. Groq itself, now inside Nvidia, plans to be among the earliest adopters. CoreWeave has deployed Spectrum-X Multiplane connecting Vera Rubin racks, and SpaceXAI says Vera CPUs will power its next generation of agentic AI. Nvidia says the racks come online later this year.

The context nobody should skip

This lands the day before Nvidia reports quarterly earnings, with analysts expecting revenue to roughly double year over year. It also lands while investors are asking harder questions about how much of Nvidia's demand is being financed by Nvidia itself, including a reported guarantee of as much as $105 billion supporting OpenAI's Ohio data-center lease.

A chip entering full production is a real, verifiable milestone. Benchmark claims from the company selling the chip are a different category. The Artificial Analysis numbers come from an independent outfit, but the configuration and the comparison point were chosen by Nvidia.

The part that will eventually matter to you has nothing to do with rack density. If this works as advertised, the agent you're using stops feeling like a slow assistant and starts feeling like something operating at conversational speed. That's a bigger product change than any benchmark chart makes it look.

Latest Andru Edwards Videos

Advertisement

Advertisement