Skip to main content
LLMgram · AI News · 2026-08-24

Nvidia Groq 3 LPX racks hit 3,400 tok/s on Gemma 4 31B with 100k-token context in Artificial Analysis test

Nvidia Groq 3 LPX racks hit 3,400 tok/s on Gemma 4 31B with 100k-token context in Artificial Analysis test

Nvidia is promoting a headline throughput result for Groq 3 LPX rack hardware tied to its Groq LPU bet. Reporting summarized on Techmeme and attributed to The Register says Nvidia claims the racks reached 3,400 tokens per second in an Artificial Analysis benchmark running Gemma 4 31B with a 100,000-token input sequence. CNBC coverage in the same packet adds that Groq racks are expected online within the year following a reported $20 billion transaction. For builders weighing long-context inference economics, the figure suggests LPUs may merit comparison with GPU serving stacks. The caveat is material: the metric is a fixed large-context throughput snapshot, not mixed-traffic latency, and adjacent coverage flags promotional framing, product-naming confusion, and community pushback on Artificial Analysis composite rankings.

Sources

Nvidia Groq 3 LPX racks hit 3,400 tok/s on Gemma 4 31B with 100k-token context in Artificial Analysis test

Nvidia Groq 3 LPX racks hit 3,400 tok/s on Gemma 4 31B with 100k-token context in Artificial Analysis test

The Register reports that Nvidia says its Groq 3 LPX racks delivered 3,400 tokens per second in an Artificial Analysis benchmark running Gemma 4 31B with a 100,000-token input sequence. The figure is a fixed large-context throughput snapshot rather than a mixed-traffic latency benchmark.

Key takeaway

A concrete LPU throughput number for 100K-token Gemma 4 31B is notable, but it is still a single marketing-adjacent datapoint rather than proof of broad inference superiority.

What happened

The Register reports, via Techmeme, that Nvidia says its Groq 3 LPX racks delivered 3,400 tokens per second in an Artificial Analysis benchmark running Gemma 4 31B with a 100,000-token input sequence.

Coverage in the packet frames the figure against Nvidia's reported multibillion-dollar Groq investment and planned rack deployments, while stressing it measures throughput on a fixed large-context workload rather than latency or mixed production traffic.

Evidence

  • Nvidia says Groq 3 LPX racks reached 3,400 tokens per second on Gemma 4 31B with a 100,000-token input in an Artificial Analysis benchmark.

    Techmeme · attributed

    Nvidia says its Groq 3 LPX racks delivered 3,400 tokens per second in an Artificial Analysis benchmark running Gemma 4 31B with a 100,000-token input sequence

  • The benchmark is a fixed large-context throughput snapshot, not a mixed-traffic latency test.

    Techmeme · attributed

    The figure is a fixed large-context throughput snapshot rather than a mixed-traffic latency benchmark.

  • Nvidia announced Groq server racks are expected to be operational within the year after a reported $20 billion transaction tied to Groq LPUs.

    CNBC AI · attributed

    Nvidia announced that Groq's server racks will be operational within the year, following a reported $20 billion transaction tied to the accelerator startup's LPUs (Linear Processing Units).

  • Nvidia denied plans to ship Groq-based LPUs to China by year-end, stating there is no China-specific LPU product in its roadmap.

    Tom's Hardware AI · attributed

    Nvidia has publicly refuted reports that it plans to ship Groq-based Linear Processing Units (LPUs) to China by year-end, stating there is no such product in its roadmap.

  • A top r/LocalLLaMA post criticizes Artificial Analysis Intelligence composite scores as misleading for production model selection.

    r/LocalLLaMA Top · attributed

    A top post on r/LocalLLaMA criticizes Artificial Analysis's 'Intelligence' benchmark as meaningless, using the release of Qwen 3.8 27B as a case study where high aggregate scores do not correlate with practical utility.

Why it matters

Inference cost and throughput at very long context lengths are now central infrastructure decisions, and this result puts LPUs on the table for that workload class even though real serving stacks rarely match a single benchmark profile.

Limits and uncertainties

The headline benchmark isolates one model size, one 100K-token context length, and one input profile, so 3,400 tok/s may not generalize to short-context or low-batch production traffic.

Nvidia is both the investor and the party promoting the result, and packet analysis flags the Groq 3 LPX naming as potentially conflated with independent Groq Inc. hardware.

Community criticism in the packet questions whether Artificial Analysis composite intelligence scores should guide production decisions without task-specific validation.

Practical implications

Operators should run benchmarks that mirror their actual request mix, batching, and latency targets before committing to LPU racks for long-context serving.

Teams should treat Artificial Analysis throughput figures and composite intelligence rankings as inputs, not substitutes, for workload-specific evaluation pipelines.

What to watch

Whether Groq racks actually come online within the year as CNBC coverage in the packet reports.

Independent verification or broader Artificial Analysis releases beyond this single Gemma 4 31B 100K-token throughput point.

Clarification of Groq 3 LPX product identity and any named customer deployments such as the Nebius reference in SiliconANGLE coverage cited on Techmeme.

Sources

LLMgram editorial selection and synthesis · @llmgram. LLMgram is not the original publisher of this information.
Continue on LLMgram: Open in AI Signal →
Original reporting: Nvidia says its Groq 3 LPX racks delivered 3,400 tokens per second in an Artificial Analysis benchmark running Gemma 4 31B with a 100,000-token input sequence (The Register)