Tag

Ai Inference

All articles tagged with #ai inference

Nvidia scales agentic AI performance with Groq 3 LPX in full production
technology1 day ago

Nvidia scales agentic AI performance with Groq 3 LPX in full production

Nvidia says its Groq 3 LPX dedicated AI inference accelerators are in full production, augmenting the Vera Rubin NVL72 platform to offload decode workloads and boost token-generation speed for complex agentic AI tasks. The system supports rack-scale deployments with up to 256 LP30 accelerators, aiming to deliver ultra-fast, real-time reasoning; Nebius is already deploying it via the Nebius Token Factory, reporting up to 3,400 tokens per second on the Gemma 4 31B model with a 100,000-token context, and SpaceX is also aligning its next-generation Vera Rubin-based AI architecture.

d-Matrix Unveils Raptor 3D-DRAM: A High-Bandwidth Path for Generative Inference
technology1 day ago

d-Matrix Unveils Raptor 3D-DRAM: A High-Bandwidth Path for Generative Inference

At Hot Chips 2026, d-Matrix unveils Raptor, a 3D-DRAM accelerator that stacks a logic die on a DRAM layer to deliver ultra-high bandwidth for generative AI inference. With a 1-Hi 32GB per card and a dense bank/channel layout, it uses stream-blocking and stream-flipping to achieve around 100 TB/s I/O and 0.37 pJ/bit, claiming ~32.6 GB/s per mm2 and 2.96 mW per GB/s versus HBM4. The system targets roughly 1,000 tokens/sec per user for frontier 3T-class models at 1M context, but faces intertwined challenges in bank-to-channel mapping, I/O power, and thermal reliability that will influence real-world viability.

HBF Emerges as a NAND-Based Bridge to Supercharge AI Inference
technology21 days ago

HBF Emerges as a NAND-Based Bridge to Supercharge AI Inference

SK hynix and SanDisk unveiled High Bandwidth Flash (HBF), a NAND-based standard meant to sit between HBM and SSDs. Stacking NAND dies to up to 512GB per configuration, HBF delivers 0.4–3 TB/s bandwidth and is designed to expand AI inference memory by storing large models that are impractical for HBM, while signaling interoperability with CPUs/GPUs via the UCIe interconnect. It is positioned to complement, not replace, HBM and to foster an ecosystem of compatible products for faster AI workloads.

AI Inference Gains a Near-Compute Memory Standard with HBF OCP Spec
technology21 days ago

AI Inference Gains a Near-Compute Memory Standard with HBF OCP Spec

SanDisk and SK hynix released the High Bandwidth Flash (HBF) technical specification through the Open Compute Project to standardize a near‑compute, high‑bandwidth memory solution for AI inference; with Google and Tenstorrent joining as consortia members, the spec defines system interfaces, electrical and packaging guidelines, reliability, and a software read/write guide, enabling HBF to coexist with High Bandwidth Memory and accelerating AI workloads. The spec is openly published to foster an early ecosystem, and Sandisk will discuss NAND’s role in AI at the FMS conference in Santa Clara.

Cerebras Delivers Breakneck LLM Speed, Yet Nvidia's CUDA Gravity Dominates
technology1 month ago

Cerebras Delivers Breakneck LLM Speed, Yet Nvidia's CUDA Gravity Dominates

NVIDIA just reported a blockbuster data-center quarter powered by CUDA, building a massive software moat, while Cerebras touts wafer-scale speed but braces for negative margins and heavy, CUDA‑centric integration that requires specialized compilation and custom engineering. Despite a $20B+ OpenAI inference deal and benchmarks showing ~21x latency advantage, the lack of broad framework support outside CUDA and the threat of OpenAI’s Jalapeño chip suggest Nvidia’s platform advantage remains hard to dethrone in the near term.

Raymond James backs Nvidia ahead of Q1 with $323 target
market-news3 months ago

Raymond James backs Nvidia ahead of Q1 with $323 target

Raymond James keeps Nvidia (NVDA) at Strong Buy with a $323 target ahead of its May 21 Q1 report, citing a growing product roadmap and AI-inference ramp. The firm also notes Nvidia’s long-term outlook of roughly $1 trillion in cumulative GPU revenue through 2027 and believes the stock remains attractive on a roughly 18x 2027 P/E, with broader Street consensus implying about 22% upside.

Nvidia Rolls Out Groq-Accelerated Inference Rack and Vera CPU Suite at GTC 2026
technology5 months ago

Nvidia Rolls Out Groq-Accelerated Inference Rack and Vera CPU Suite at GTC 2026

Nvidia announced at GTC 2026 a Groq-based 3 LPX inference rack (256 LPUs, liquid-cooled, Spectrum-X interconnect) paired with the Vera Rubin NVL72 rack, plus a Vera CPU rack and BlueField-4 STX storage rack, all aimed at accelerating trillion-parameter AI models and agentic AI workloads. Rubin CPX products were put on hold to focus on LPX, with availability in H2 2026 and OEMs like Dell, HPE and Oracle Cloud.

Maia 200 Pushes Cloud AI In-House, But Nvidia Keeps the Data Center Edge
technology7 months ago

Maia 200 Pushes Cloud AI In-House, But Nvidia Keeps the Data Center Edge

Microsoft’s Maia 200 is an in‑house AI inference accelerator for Azure that claims strong performance per dollar and will power OpenAI models, signaling rising cloud‑provider pressure on Nvidia. While Maia 200 underscores a shift toward custom silicon, Nvidia still leads the data‑center AI market with its broad GPU ecosystem and software stack, and though cloud‑provider alternatives may erode pricing power over time, a rapid disruption to Nvidia’s position appears unlikely, even as valuations remain rich given AI growth.

Nvidia's Strategic Partnership with Groq Boosts AI Chip Competition and Stock
business8 months ago

Nvidia's Strategic Partnership with Groq Boosts AI Chip Competition and Stock

Nvidia's strategic licensing agreement with AI startup Groq, including key personnel hires, aims to strengthen its position in AI inference technology, signaling a shift from training to inference workloads and potentially expanding Nvidia's market dominance. The deal, which keeps Groq independent, is viewed positively by analysts as a move to address market share concerns and diversify Nvidia's AI offerings.