Tag

Inference

All articles tagged with #inference

AMD bets on silicon-etched AI models with Taalas acquisition
technology18 days ago

AMD bets on silicon-etched AI models with Taalas acquisition

AMD has acquired Toronto-based AI-chip startup Taalas, which bakes model weights directly into silicon to create model-specific MSICs and boost inference speeds. Early tests of Taalas’ HC1 chip reached about 16,960 tokens per second (roughly 48x faster than Nvidia GPUs), with HC2 targeting 20 billion parameters. AMD plans a disaggregated architecture pairing Instinct GPUs with Taalas accelerators, potentially lowering per-token costs but requiring model re-spins when weights change. The deal is expected to close in Q4, pending regulatory approval.

AI Coding Agents Put Nvidia's CUDA Edge to the Test
technology22 days ago

AI Coding Agents Put Nvidia's CUDA Edge to the Test

AI coding agents are increasingly able to reproduce CUDA-like software, threatening Nvidia's long-standing software moat. While startups and cloud providers explore cross-chip compatibility and open approaches, Nvidia argues that its tightly integrated hardware-software stack and robust verification tools still create a durable advantage, suggesting the battle may hinge on how quickly tooling and inference workloads evolve rather than a collapse of CUDA.

SambaNova valued at $11B after $1B funding to accelerate on-prem AI inference
technology1 month ago

SambaNova valued at $11B after $1B funding to accelerate on-prem AI inference

SambaNova raised $1 billion in fresh financing led by General Atlantic (with Seligman Ventures, T. Rowe Price and Capital Group), boosting its valuation to $11 billion and fueling the rollout of its SN50 inference chips for on‑premise deployments. JPMorgan has adopted the company's systems for demanding enterprise AI workloads, while the startup—which earlier drew funding from Intel—continues to pursue a U.S. IPO around 2027 amid strong investor interest in AI-chip companies challenging Nvidia.

OpenAI and Broadcom unveil Jalapeño, a data-center chip for scalable LLM inference
technology2 months ago

OpenAI and Broadcom unveil Jalapeño, a data-center chip for scalable LLM inference

OpenAI and Broadcom introduced Jalapeño, a purpose-built ASIC designed from scratch for large-language-model inference in data centers, with early testing claiming substantially better performance per watt; development took nine months and is part of a broader effort to own more of the AI stack and reduce reliance on Nvidia, with deployments planned by year-end as the silicon race heats up.

Intel bets on a cheaper AI inference GPU to shake up data-center chips this year
technology2 months ago

Intel bets on a cheaper AI inference GPU to shake up data-center chips this year

Intel plans to ship its Crescent Island AI data-centre GPU by year-end to accelerate inference, using cheaper LPDDR5 memory and air cooling to cut cost vs Nvidia/AMD. Led by Kevork Kechichian, it marks Intel’s first major AI infrastructure push under CEO Lip-Bu Tan, with limited initial shipments after an 18-month development. The chip targets cheaper memory and cooling, aims to compete on price and power, and Intel is pursuing in-house foundry plans with potential China sales under export controls as it rebuilds its AI hardware business.

Memory Over Speed: The AI Inference Shift
technology3 months ago

Memory Over Speed: The AI Inference Shift

Ben Thompson argues that the AI compute boom is moving from GPU-dominated training to memory-centric, agentic-inference architectures; Cerebras’ wafer-scale chips offer extraordinary on-chip memory and bandwidth for fast answer inference but face cost and scalability limits, while the long-term potential lies in memory hierarchies that support autonomous agentic work, potentially reducing Nvidia’s dominance and reconfiguring compute across training, inference, and even space data centers.

Nvidia Unveils Groq-Enhanced Inference to Defend AI Chip Lead
technology5 months ago

Nvidia Unveils Groq-Enhanced Inference to Defend AI Chip Lead

At its GTC conference, Nvidia unveiled a product that pairs its chips with Groq’s acceleration tech to boost AI inference speed and cut costs, a move aimed at defending its dominant hardware position as rivals advance. The announcement follows Nvidia’s $20 billion Groq licensing deal and includes NemoClaw to help software companies deploy AI agents, all while supply-chain and manufacturing constraints shape growth prospects for its Rubin and Blackwell line.

Nvidia-OpenAI $100B plan fizzles into non-binding talks
technology6 months ago

Nvidia-OpenAI $100B plan fizzles into non-binding talks

A September 2025 letter of intent for Nvidia to invest up to $100 billion in OpenAI’s AI infrastructure has not materialized; Nvidia’s Jensen Huang says the figure was never a commitment, and Reuters reports OpenAI has been seeking alternatives and citing Nvidia chip speed concerns for inference. OpenAI has since struck deals with Cerebras, Groq, AMD, and Broadcom to diversify compute, while Nvidia emphasizes a large future investment but not at that scale. The news triggered a stock dip for Nvidia and highlighted questions about timing and strategic fit.

OpenAI weighs chip alternatives after Nvidia inference gaps
technology6 months ago

OpenAI weighs chip alternatives after Nvidia inference gaps

OpenAI is reportedly seeking alternatives to Nvidia GPUs due to dissatisfaction with inference performance, citing eight sources. The move follows reports that Nvidia’s plan to invest up to $100 billion in OpenAI has stalled. OpenAI has previously struck deals with AMD and Broadcom to develop custom AI accelerators, signaling a push to diversify hardware sources even as Nvidia remains a major partner.