Tag

Edge Computing

All articles tagged with #edge computing

Nvidia PAIR Turns Your Idle Computers Into a Private Home AI Network
technology1 month ago

Nvidia PAIR Turns Your Idle Computers Into a Private Home AI Network

Nvidia unveiled PAIR (Personal AI Router), free/open-source software that turns idle home devices—PCs, laptops, Macs—into a coordinated, private AI cluster. It acts as an intelligent traffic controller, breaking tasks into sub-tasks and routing them to capable machines across the local network for parallel processing. It doesn’t pool VRAM or split a single model; instead, it assigns complete tasks to individual devices, supporting cross‑platform hardware (GeForce GPUs, Apple silicon, DGX Spark) and Windows/macOS/Linux, with no cloud subscription and even offline operation after models are downloaded. The setup requires about 8GB RAM and 20GB+ disk per machine and aims to let households run local AI securely and privately, at the edge.

NVIDIA PAIR Turns Idle Home PCs Into a Private AI Cluster, Cutting Cloud Inference Bills
technology1 month ago

NVIDIA PAIR Turns Idle Home PCs Into a Private AI Cluster, Cutting Cloud Inference Bills

NVIDIA unveils PAIR, a software router that distributes AI inference across devices on a home network to create a private, cloud-free AI cluster. It auto-discovers machines, uses secure MTLS, and proxies via Ollama and LM Studio, with cross-platform support (Windows, Linux, macOS) and an open-source Apache 2.0 beta release. By routing idle GPU/CPU workloads to available nodes and avoiding apps like gaming on busy machines, PAIR aims to reduce cloud API usage—statements suggest substantial savings (e.g., around $1,200/month in cloud credits for certain workloads)—while keeping data and prompts local. It does not merge VRAM or tensor-split models; instead it uses a queue-based scheduler for multi-agent workflows, multi-tasking, and system offload, with planned future enhancements.

FreeToken Lets 753B GLM-5.2 Run on a Single Desktop GPU
technology1 month ago

FreeToken Lets 753B GLM-5.2 Run on a Single Desktop GPU

Researchers unveil FreeToken, an edge-native Mixture-of-Experts serving engine that maps model state and computation to a user’s hardware, enabling massive frontier models like 753B GLM-5.2 to run on a single workstation GPU. Demonstrations show 35B on an 8 GB laptop GPU, 284B on a gaming desktop, and the full 753B on a single workstation GPU, with OpenAI/Anthropic-compatible endpoints and a Windows/Linux desktop app. FreeToken achieves 1.5–2.3× decode throughput over llama.cpp, Ollama, and KTransformers, thanks to bandwidth-aware scheduling, semantic-aware caching, and elastic memory management that keep results exact and avoid heavy CPU offloads. It's Apache-2.0, pip-installable (freetoken), and designed for solo developers, SMBs, and regulated workloads where data stays on-device; targeted use cases include private code analysis, offline contract review, synthetic data, and local agent workloads.

Brain-inspired memtransistor chip promises ultra-efficient, edge-friendly AI
technology1 month ago

Brain-inspired memtransistor chip promises ultra-efficient, edge-friendly AI

Researchers unveiled a cerebellum-inspired memtransistor chip that merges memory and processing to power neuromorphic AI. In ECG simulations, the device detected arrhythmias with 98% accuracy, ran about twice as fast as a transformer, and used roughly 10,000 times fewer computations, signaling potential for fast, energy-efficient edge AI and reduced reliance on data centers for responsive applications such as health monitoring, autonomous cars, and robotics. Built from molybdenum disulfide, the memtransistor forms the core of the output layer of a spiking neural network, offering a path toward low-power hardware that reacts to unexpected events instead of processing all data.

AI on a Budget: Tech Giants Pivot from Token Frenzy to Cheaper Models
artificial-intelligence4 months ago

AI on a Budget: Tech Giants Pivot from Token Frenzy to Cheaper Models

Facing soaring token costs, Big Tech is rethinking AI adoption: Amazon and Uber cap usage, OpenAI's Sam Altman calls token usage a huge issue, and firms pivot toward leaner edge‑computing models to cut token bills; GitHub's pay‑per‑token plan struggles, while Microsoft and Google push edge solutions to reduce cloud reliance, even as data‑center energy and water concerns linger.

RTX Spark, Solara, and the Cloud-First AI Reboot
technology4 months ago

RTX Spark, Solara, and the Cloud-First AI Reboot

Ben Thompson surveys Nvidia’s RTX Spark AI PC, Project Solara, and Microsoft’s in-house MAI initiative, arguing the RTX Spark’s local inference and CPU/GPU balance may not meet the needs of an “agentic” AI era, while Microsoft pushes a cloud-centric, enterprise-ready AI stack that enables custom, data-owned models and a platform-spanning agent paradigm that could reshape devices and workflows.

Apple Pushes On-Device AI Ahead of WWDC, Reducing Cloud Reliance
technology4 months ago

Apple Pushes On-Device AI Ahead of WWDC, Reducing Cloud Reliance

Apple plans to foreground AI that runs on iPhones and other devices at WWDC, leveraging its 15 years of silicon design to speed edge AI. However, many queries will still require cloud processing for complexity, with Siri tasks routing to Google Cloud’s Gemini and Nvidia privacy tech used in that setting, signaling a hybrid approach that boosts local AI while maintaining cloud-backed capabilities.

Homes as Micro Data Centers: A Creeping AI Infrastructure Trend
technology5 months ago

Homes as Micro Data Centers: A Creeping AI Infrastructure Trend

The piece explores the notion of turning homes into micro data centers to support AI workloads, highlighting pilots like Span with Nvidia and PulteGroup and heat-reuse experiments, while noting regulatory pushback on new centers and public concern over power use. While some workloads could work in a residential setting, experts say power, cooling, connectivity, security, and regulatory hurdles make home-centered data centers likely a niche rather than a replacement for hyperscale facilities.

technology5 months ago

Edge AI Push Could Lift Qualcomm to $340, Trefis Analysis

Trefis argues Qualcomm is transitioning from a handset component supplier to an edge‑AI compute platform, leveraging Snapdragon, automotive design-wins and on‑device inferencing. Despite near‑term headwinds like Apple’s modem shift and memory shortages, a growth path driven by AI, automotive, and on‑device compute could push the stock toward $340; with revenue seen rising to about $65B by 2029 and EPS near $17, a ~20x forward multiple implies the target, supported by strong cash flow and buybacks.

NVIDIA Takes AI to Orbit with Space Computing Platform
technology6 months ago

NVIDIA Takes AI to Orbit with Space Computing Platform

NVIDIA announced Space Computing, bringing data-center‑class AI to orbital data centers and space-edge environments via the Space‑1 Vera Rubin Module, IGX Thor, and Jetson Orin. Partners such as Aetherflux, Axiom Space, Kepler, Planet, Sophia Space and Starcloud will power next‑gen space missions with on‑board AI, real‑time analytics and autonomous operations. The lineup also includes ground‑processing acceleration with RTX PRO 6000 Blackwell Server Edition, while Space‑1 Rubin promises up to ~25x more AI compute for space inferencing versus the H100; IGX Thor and Jetson Orin are available now, with Space‑1 Rubin Module coming later.

Autonomous cytopathology hits clinical-grade accuracy with real-time 3D slide tomography
technology7 months ago

Autonomous cytopathology hits clinical-grade accuracy with real-time 3D slide tomography

Researchers report a clinical-grade autonomous cytopathology platform that uses real-time 3D whole-slide tomography with edge computing and CMD-based population analysis to automatically detect and classify cervical cells, achieving single-cell AUC over 0.99 and slide-level AUCs up to 0.97 across 1,124 samples, enabling autonomous triage cytology and objective, scalable diagnostics.

Raspberry Pi's AI HAT+ 2 brings 8GB RAM to edge AI, but power and price temper expectations
technology8 months ago

Raspberry Pi's AI HAT+ 2 brings 8GB RAM to edge AI, but power and price temper expectations

Raspberry Pi's AI HAT+ 2 adds 8GB RAM and a Hailo 10H chip to run small generative AI models locally on the Pi 5 (Llama 3.2, Qwen) with training support, at a $130 price, but independent tests suggest the Pi 5 with 8GB RAM often outperforms the add-on due to its higher power budget, making the upgrade less compelling than simply using a higher-RAM Pi or the cheaper AI HAT+ variant.