Nvidia Groq 3 LPX Enters Full Production With Nebius as First Customer
Reporting circulated that Nvidia has positioned Groq-based Linear Processing Unit racks for datacenter inference, with SiliconANGLE coverage cited on Techmeme claiming an accelerator labeled Groq 3 LPX reached full production and Nebius signed as first customer while SpaceX would deploy Vera CPUs. Separately, Nvidia said Groq racks should be online this year after a reported twenty-billion-dollar deal, and claimed three thousand four hundred tokens per second on Gemma 4 31B with a one-hundred-thousand-token context in an Artificial Analysis benchmark. Neocloud comparisons place Nebius among major GPU operators after Groq licensed LPU technology to Nvidia. Nvidia also denied plans for China-specific LPU shipments. The Groq 3 LPX production label collides with Groq Inc.'s existing inference brand, and benchmark figures reflect one promotional workload rather than broad serving proof.
Nvidia Groq 3 LPX Enters Full Production With Nebius as First Customer
Nvidia says its inference accelerator Groq 3 LPX has entered full production and Nebius has signed on as the first customer, and SpaceX will deploy Vera CPUs. Mike Wheatley / SiliconANGLE reports the dedicated artificial intelligence inference accelerator Groq 3 LPX has now entered full production.
Key takeaway
Nvidia is signaling Groq LPU racks moving toward production with Nebius as anchor customer, but the Groq 3 LPX naming clash with Groq Inc. chips means the report needs official confirmation before acting.
What happened
Per Mike Wheatley reporting in SiliconANGLE and cited on Techmeme, Nvidia says its dedicated artificial intelligence inference accelerator Groq 3 LPX has entered full production, Nebius has signed on as the first customer, and SpaceX will deploy Vera CPUs.
CNBC coverage states Nvidia said Groq racks will be online this year following a reported twenty-billion-dollar purchase tied to Groq LPUs, and The Register reports Nvidia claimed Groq 3 LPX racks delivered 3,400 tokens per second in an Artificial Analysis benchmark running Gemma 4 31B with a 100,000-token input sequence.
Evidence
Nvidia says Groq 3 LPX entered full production with Nebius as first customer
Techmeme · attributed
Nvidia says its inference accelerator Groq 3 LPX has entered full production and Nebius has signed on as the first customer, and SpaceX will deploy Vera CPUs
Nvidia claimed 3,400 tokens per second on a 100,000-token Gemma 4 31B benchmark
Techmeme · attributed
Nvidia says its Groq 3 LPX racks delivered 3,400 tokens per second in an Artificial Analysis benchmark running Gemma 4 31B with a 100,000-token input sequence
Nvidia said Groq racks will be online this year after a reported $20 billion purchase
CNBC AI · attributed
Nvidia says Groq racks will be online this year following $20 billion purchase
Groq rebuilt as an inference cloud after licensing LPU technology to NVIDIA
MarkTechPost · attributed
Groq rebuilt itself as an inference cloud after licensing its LPU technology to NVIDIA
Nvidia denied a China-specific LPU product roadmap
Tom's Hardware AI · attributed
Nvidia denies report it will ship Groq-based LPUs to China by year-end — says there is 'no China-specific LPU product in our roadmap'
The Groq 3 LPX label may conflate Nvidia reporting with Groq Inc. inference chips
Techmeme · attributed
Nvidia's naming of its inference accelerator 'Groq 3 LPX' conflates it with the independent Groq startup, and the claim warrants scrutiny because Groq-branded inference chips already exist from a different company.
Why it matters
If Groq LPU racks reach commercial scale, operators gain a specialized inference option for long-context serving beyond GPU clusters, but product identity and workload-specific benchmarks must be validated before infrastructure commitments.
Limits and uncertainties
The Groq 3 LPX production label may be a reporting error because Groq Inc. already sells inference chips under the Groq brand.
SiliconANGLE and Techmeme excerpts are truncated and lack verified performance, capacity, or deployment specifics.
The 3,400 tokens per second figure reflects one model, one context length, and one benchmark profile rather than mixed production traffic.
CNBC and Tom's Hardware excerpts in the packet are brief and do not fully document deal terms or the scope of Nvidia's denial.
Practical implications
Treat Nebius as an early signal for LPU rack availability but confirm product naming and specs through Nvidia official channels before engineering plans.
Request benchmarks that match your actual request patterns, batch sizes, and context lengths rather than relying on a single Artificial Analysis datapoint.
Factor Groq's pivot to an inference cloud and Nebius SEC reporting into neocloud vendor comparisons alongside headline GPU pricing.
What to watch
Nvidia official statements clarifying Groq LPU rack naming, production status, and Vera CPU deployment timelines.
Nebius public disclosures or service announcements confirming Groq LPU rack capacity and customer availability.
Independent benchmark releases covering short-context and mixed-traffic inference beyond the Gemma 4 31B 100,000-token test.
Original reporting: Nvidia says its inference accelerator Groq 3 LPX has entered full production and Nebius has signed on as the first customer, and SpaceX will deploy Vera CPUs (Mike Wheatley/SiliconANGLE)