Nvidia’s Groq-3 LPUs hit 3,400 tok/s on Gemma-4-31B, signaling a bold AI-inference bet

1 min read
Source: The Register
Nvidia’s Groq-3 LPUs hit 3,400 tok/s on Gemma-4-31B, signaling a bold AI-inference bet
Photo: The Register
TL;DR Summary

Nvidia’s Groq 3 LPU-based LPX racks achieved 3,400 tokens per second on the Gemma 4 31B model in an independent Artificial Analysis benchmark, claimed to be about 4x faster than Cerebras for this scenario. The design relies on SRAM-heavy LPUs with high memory bandwidth and distributes the model across multiple chips via Ethernet, while GPUs handle the prefill phase and LPUs handle the decode phase in a heterogeneous inference setup. Nebius is among the first customers. While the result shows strong throughput, the 31B model is relatively small and dense, and scaling to larger MoE models remains uncertain, with real-world performance depending on model size, distribution, and upcoming accelerator generations.

Share this article

Reading Insights

Total Reads

0

Unique Readers

8

Time Saved

6 min

vs 7 min read

Condensed

92%

1,379110 words

Want the full story? Read the original article

Read on The Register