Nvidia scales agentic AI performance with Groq 3 LPX in full production

TL;DR
Nvidia says its Groq 3 LPX dedicated AI inference accelerators are in full production, augmenting the Vera Rubin NVL72 platform to offload decode workloads and boost token-generation speed for complex agentic AI tasks. The system supports rack-scale deployments with up to 256 LP30 accelerators, aiming to deliver ultra-fast, real-time reasoning; Nebius is already deploying it via the Nebius Token Factory, reporting up to 3,400 tokens per second on the Gemma 4 31B model with a 100,000-token context, and SpaceX is also aligning its next-generation Vera Rubin-based AI architecture.
- Nvidia's dedicated inference accelerator Groq 3 LPX enters full production to supercharge AI agents SiliconANGLE
- Nvidia says Groq racks will be online this year following $20 billion purchase CNBC
- NVIDIA Groq 3 LPX Now in Full Production With World-Class Speed for Agentic AI NVIDIA Newsroom
- What Nvidia's first Groq 3 LPU benchmarks tell us about its $20B gamble The Register
- Nvidia’s $20 Billion Groq Bet Is Going Live Before the End of 2026. Here’s What It Means for Investors. The Motley Fool
Want the full story? Read the original reporting
Read on SiliconANGLE