NVIDIA Puts Groq 3 LPX Into Full Production, Aiming Squarely at Agent Latency
NVIDIA's inference accelerator hit 3,400 output tokens per second on Gemma 4 31B at 100,000-token context. Nebius is the first cloud to adopt it.

NVIDIA announced at Hot Chips that Groq 3 LPX, its interactive AI inference accelerator, is now in full production. It extends the Vera Rubin platform, and it is built for one specific bottleneck: how fast tokens come out.
Why token generation rate is the constraint
An agentic system does not make one model call. It makes hundreds or thousands of inference steps, generating tokens across all of them, before it finishes a task. Every step adds latency, and latency compounds. That is the difference between an agent that completes a coding task in minutes and one that takes hours.
Vera Rubin NVL72 systems handle training and inference for AI factories generally. Groq 3 LPX extends their inference side by raising the token generation rate, which is the number that governs how responsive an agent feels on context-heavy work.
The benchmark numbers
In Artificial Analysis benchmarking, Groq 3 LPX recorded 3,400 output tokens per second running Gemma 4 31B, an open-source agentic model, at 100,000-token context. NVIDIA says that is the fastest performance ever recorded for that model.
The claim it attaches to that: 4x faster responsiveness for agents and latency-sensitive workloads than the nearest alternative platform, and agentic coding tasks that finish in minutes rather than hours.
Benchmark figures always come from the vendor announcing them. The 100,000-token context is the part worth noting, because short-context throughput numbers rarely survive contact with real agent workloads.
What Jensen Huang said
Huang framed inference as the growth engine of AI, describing Vera Rubin as extending the Grace Blackwell vision with workload-optimised AI factory configurations for agentic systems, with LPX handling ultrafast token generation.
The two-problem framing
NVIDIA's own description of the design problem is the clearest part of the announcement. Agentic AI creates two distinct computing challenges: processing enormous amounts of context, and generating tokens at extremely low latency. Those pull hardware in different directions. Groq 3 LPX is purpose-built for the second while riding on Vera Rubin for the first.
Nebius is the first AI cloud to adopt it.
Source: NVIDIA Groq 3 LPX Now in Full Production With World-Class Speed for Agentic AI — NVIDIA Newsroom