NVIDIA says Groq 3 LPX enters full production

NVIDIA said Aug. 24 that its Groq 3 LPX inference system has entered full production alongside the Vera Rubin NVL72 platform. In a 100,000-token Gemma 4 31B benchmark, NVIDIA reported about 3,400 output tokens per second; the result applies to a named vendor configuration, not every deployment.
NVIDIA said on August 24 that its Groq 3 LPX inference system has entered full production alongside the Vera Rubin NVL72 platform. In an Artificial Analysis 100,000-token benchmark on Gemma 4 31B, NVIDIA reported about 3,400 output tokens per second.
A production claim with a specific benchmark
The announcement concerns an NVIDIA configuration for long-context inference, rather than a general performance result for every model or deployment. NVIDIA's technical account gives a more precise median result of 3,431 output tokens per second in the cited benchmark and says the comparison used a 100,000-token input context.
The setup combines the LPX system with Vera Rubin NVL72. NVIDIA describes the LPX rack as 256 interconnected local processing units, designed to prioritize deterministic token generation and low latency for inference workloads. Those architectural claims come from the vendor and should be read as product specifications rather than an independent capacity study.
What the result does and does not show
A 100,000-token test is relevant to workloads that keep substantial context in play, such as code agents and retrieval-heavy systems. Still, the reported token rate is tied to a named model, a particular benchmark, and NVIDIA-operated hardware. It does not establish the latency, cost, reliability, or throughput that another operator will see with a different model, batch size, context length, or serving stack.
Groq said it intends to bring the platform to market through its inference cloud, while StorageReview separately reported NVIDIA's full-production announcement and the same headline benchmark. The immediate reader takeaway is therefore narrower than the marketing language: NVIDIA is presenting a production-ready long-context inference configuration with a high benchmark result, but buyer evaluation still requires workload-specific tests and availability details.
Key Points
- 1NVIDIA says Groq 3 LPX entered full production on August 24 as part of the Vera Rubin NVL72 platform.
- 2NVIDIA reported about 3,400 output tokens per second for Gemma 4 31B with a 100,000-token input in an Artificial Analysis benchmark.
- 3The vendor result does not substitute for workload-specific tests of latency, cost, availability, and reliability.
Scoring Rationale
NVIDIA's full-production claim and long-context benchmark are material infrastructure signals for teams evaluating high-interactivity agent workloads, while the disclosed measurement remains vendor-specific and does not establish general deployment performance.
Sources
Primary source and supporting public references used for this report.
Practice interview problems based on real data
1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with.
Try 250 free problems

