Skip to main content
LLMgram · AI News · 2026-08-24

Nvidia Groq 3 LPX Enters Full Production With Nebius as First Customer

Nvidia Groq 3 LPX Enters Full Production With Nebius as First Customer

Reporting circulated that Nvidia has positioned Groq-based Linear Processing Unit racks for datacenter inference, with SiliconANGLE coverage cited on Techmeme claiming an accelerator labeled Groq 3 LPX reached full production and Nebius signed as first customer while SpaceX would deploy Vera CPUs. Separately, Nvidia said Groq racks should be online this year after a reported twenty-billion-dollar deal, and claimed three thousand four hundred tokens per second on Gemma 4 31B with a one-hundred-thousand-token context in an Artificial Analysis benchmark. Neocloud comparisons place Nebius among major GPU operators after Groq licensed LPU technology to Nvidia. Nvidia also denied plans for China-specific LPU shipments. The Groq 3 LPX production label collides with Groq Inc.'s existing inference brand, and benchmark figures reflect one promotional workload rather than broad serving proof.

Sources

Nvidia Groq 3 LPX Enters Full Production With Nebius as First Customer

Nvidia Groq 3 LPX Enters Full Production With Nebius as First Customer

Nvidia says its inference accelerator Groq 3 LPX has entered full production and Nebius has signed on as the first customer, and SpaceX will deploy Vera CPUs. Mike Wheatley / SiliconANGLE reports the dedicated artificial intelligence inference accelerator Groq 3 LPX has now entered full production.

Key takeaway

Nvidia is signaling Groq LPU racks moving toward production with Nebius as anchor customer, but the Groq 3 LPX naming clash with Groq Inc. chips means the report needs official confirmation before acting.

What happened

Per Mike Wheatley reporting in SiliconANGLE and cited on Techmeme, Nvidia says its dedicated artificial intelligence inference accelerator Groq 3 LPX has entered full production, Nebius has signed on as the first customer, and SpaceX will deploy Vera CPUs.

CNBC coverage states Nvidia said Groq racks will be online this year following a reported twenty-billion-dollar purchase tied to Groq LPUs, and The Register reports Nvidia claimed Groq 3 LPX racks delivered 3,400 tokens per second in an Artificial Analysis benchmark running Gemma 4 31B with a 100,000-token input sequence.

Evidence

  • Nvidia says Groq 3 LPX entered full production with Nebius as first customer

    Techmeme · attributed

    Nvidia says its inference accelerator Groq 3 LPX has entered full production and Nebius has signed on as the first customer, and SpaceX will deploy Vera CPUs

  • Nvidia claimed 3,400 tokens per second on a 100,000-token Gemma 4 31B benchmark

    Techmeme · attributed

    Nvidia says its Groq 3 LPX racks delivered 3,400 tokens per second in an Artificial Analysis benchmark running Gemma 4 31B with a 100,000-token input sequence

  • Nvidia said Groq racks will be online this year after a reported $20 billion purchase

    CNBC AI · attributed

    Nvidia says Groq racks will be online this year following $20 billion purchase

  • Groq rebuilt as an inference cloud after licensing LPU technology to NVIDIA

    MarkTechPost · attributed

    Groq rebuilt itself as an inference cloud after licensing its LPU technology to NVIDIA

  • Nvidia denied a China-specific LPU product roadmap

    Tom's Hardware AI · attributed

    Nvidia denies report it will ship Groq-based LPUs to China by year-end — says there is 'no China-specific LPU product in our roadmap'

  • The Groq 3 LPX label may conflate Nvidia reporting with Groq Inc. inference chips

    Techmeme · attributed

    Nvidia's naming of its inference accelerator 'Groq 3 LPX' conflates it with the independent Groq startup, and the claim warrants scrutiny because Groq-branded inference chips already exist from a different company.

Why it matters

If Groq LPU racks reach commercial scale, operators gain a specialized inference option for long-context serving beyond GPU clusters, but product identity and workload-specific benchmarks must be validated before infrastructure commitments.

Limits and uncertainties

The Groq 3 LPX production label may be a reporting error because Groq Inc. already sells inference chips under the Groq brand.

SiliconANGLE and Techmeme excerpts are truncated and lack verified performance, capacity, or deployment specifics.

The 3,400 tokens per second figure reflects one model, one context length, and one benchmark profile rather than mixed production traffic.

CNBC and Tom's Hardware excerpts in the packet are brief and do not fully document deal terms or the scope of Nvidia's denial.

Practical implications

Treat Nebius as an early signal for LPU rack availability but confirm product naming and specs through Nvidia official channels before engineering plans.

Request benchmarks that match your actual request patterns, batch sizes, and context lengths rather than relying on a single Artificial Analysis datapoint.

Factor Groq's pivot to an inference cloud and Nebius SEC reporting into neocloud vendor comparisons alongside headline GPU pricing.

What to watch

Nvidia official statements clarifying Groq LPU rack naming, production status, and Vera CPU deployment timelines.

Nebius public disclosures or service announcements confirming Groq LPU rack capacity and customer availability.

Independent benchmark releases covering short-context and mixed-traffic inference beyond the Gemma 4 31B 100,000-token test.

Sources

LLMgram editorial selection and synthesis · @llmgram. LLMgram is not the original publisher of this information.
Continue on LLMgram: Open in AI Signal →
Original reporting: Nvidia says its inference accelerator Groq 3 LPX has entered full production and Nebius has signed on as the first customer, and SpaceX will deploy Vera CPUs (Mike Wheatley/SiliconANGLE)