Nvidia’s Groq-3 LPUs hit 3,400 tok/s on Gemma-4-31B, signaling a bold AI-inference bet

Nvidia’s Groq 3 LPU-based LPX racks achieved 3,400 tokens per second on the Gemma 4 31B model in an independent Artificial Analysis benchmark, claimed to be about 4x faster than Cerebras for this scenario. The design relies on SRAM-heavy LPUs with high memory bandwidth and distributes the model across multiple chips via Ethernet, while GPUs handle the prefill phase and LPUs handle the decode phase in a heterogeneous inference setup. Nebius is among the first customers. While the result shows strong throughput, the 31B model is relatively small and dense, and scaling to larger MoE models remains uncertain, with real-world performance depending on model size, distribution, and upcoming accelerator generations.
- What Nvidia's first Groq 3 LPU benchmarks tell us about its $20B gamble The Register
- Nvidia says Groq racks will be online this year following $20 billion purchase CNBC
- How NVIDIA Groq 3 LPX Unlocks Ultrafast Interactivity at Long Context on NVIDIA Vera Rubin | NVIDIA Technical Blog NVIDIA Developer
- Nvidia's dedicated inference accelerator Groq 3 LPX enters full production to supercharge AI agents SiliconANGLE
- Nvidia's AI inference chip from its $20 billion Groq deal enters full production qz.com
Reading Insights
0
8
6 min
vs 7 min read
92%
1,379 → 110 words
Want the full story? Read the original article
Read on The Register