Tag

Edge Computing

All articles tagged with #edge computing

FreeToken Lets 753B GLM-5.2 Run on a Single Desktop GPU
technology1 day ago

FreeToken Lets 753B GLM-5.2 Run on a Single Desktop GPU

Researchers unveil FreeToken, an edge-native Mixture-of-Experts serving engine that maps model state and computation to a user’s hardware, enabling massive frontier models like 753B GLM-5.2 to run on a single workstation GPU. Demonstrations show 35B on an 8 GB laptop GPU, 284B on a gaming desktop, and the full 753B on a single workstation GPU, with OpenAI/Anthropic-compatible endpoints and a Windows/Linux desktop app. FreeToken achieves 1.5–2.3× decode throughput over llama.cpp, Ollama, and KTransformers, thanks to bandwidth-aware scheduling, semantic-aware caching, and elastic memory management that keep results exact and avoid heavy CPU offloads. It's Apache-2.0, pip-installable (freetoken), and designed for solo developers, SMBs, and regulated workloads where data stays on-device; targeted use cases include private code analysis, offline contract review, synthetic data, and local agent workloads.

Brain-inspired memtransistor chip promises ultra-efficient, edge-friendly AI
technology13 days ago

Brain-inspired memtransistor chip promises ultra-efficient, edge-friendly AI

Researchers unveiled a cerebellum-inspired memtransistor chip that merges memory and processing to power neuromorphic AI. In ECG simulations, the device detected arrhythmias with 98% accuracy, ran about twice as fast as a transformer, and used roughly 10,000 times fewer computations, signaling potential for fast, energy-efficient edge AI and reduced reliance on data centers for responsive applications such as health monitoring, autonomous cars, and robotics. Built from molybdenum disulfide, the memtransistor forms the core of the output layer of a spiking neural network, offering a path toward low-power hardware that reacts to unexpected events instead of processing all data.

AI on a Budget: Tech Giants Pivot from Token Frenzy to Cheaper Models
artificial-intelligence2 months ago

AI on a Budget: Tech Giants Pivot from Token Frenzy to Cheaper Models

Facing soaring token costs, Big Tech is rethinking AI adoption: Amazon and Uber cap usage, OpenAI's Sam Altman calls token usage a huge issue, and firms pivot toward leaner edge‑computing models to cut token bills; GitHub's pay‑per‑token plan struggles, while Microsoft and Google push edge solutions to reduce cloud reliance, even as data‑center energy and water concerns linger.

RTX Spark, Solara, and the Cloud-First AI Reboot
technology2 months ago

RTX Spark, Solara, and the Cloud-First AI Reboot

Ben Thompson surveys Nvidia’s RTX Spark AI PC, Project Solara, and Microsoft’s in-house MAI initiative, arguing the RTX Spark’s local inference and CPU/GPU balance may not meet the needs of an “agentic” AI era, while Microsoft pushes a cloud-centric, enterprise-ready AI stack that enables custom, data-owned models and a platform-spanning agent paradigm that could reshape devices and workflows.

Apple Pushes On-Device AI Ahead of WWDC, Reducing Cloud Reliance
technology2 months ago

Apple Pushes On-Device AI Ahead of WWDC, Reducing Cloud Reliance

Apple plans to foreground AI that runs on iPhones and other devices at WWDC, leveraging its 15 years of silicon design to speed edge AI. However, many queries will still require cloud processing for complexity, with Siri tasks routing to Google Cloud’s Gemini and Nvidia privacy tech used in that setting, signaling a hybrid approach that boosts local AI while maintaining cloud-backed capabilities.

Homes as Micro Data Centers: A Creeping AI Infrastructure Trend
technology3 months ago

Homes as Micro Data Centers: A Creeping AI Infrastructure Trend

The piece explores the notion of turning homes into micro data centers to support AI workloads, highlighting pilots like Span with Nvidia and PulteGroup and heat-reuse experiments, while noting regulatory pushback on new centers and public concern over power use. While some workloads could work in a residential setting, experts say power, cooling, connectivity, security, and regulatory hurdles make home-centered data centers likely a niche rather than a replacement for hyperscale facilities.

technology3 months ago

Edge AI Push Could Lift Qualcomm to $340, Trefis Analysis

Trefis argues Qualcomm is transitioning from a handset component supplier to an edge‑AI compute platform, leveraging Snapdragon, automotive design-wins and on‑device inferencing. Despite near‑term headwinds like Apple’s modem shift and memory shortages, a growth path driven by AI, automotive, and on‑device compute could push the stock toward $340; with revenue seen rising to about $65B by 2029 and EPS near $17, a ~20x forward multiple implies the target, supported by strong cash flow and buybacks.

NVIDIA Takes AI to Orbit with Space Computing Platform
technology5 months ago

NVIDIA Takes AI to Orbit with Space Computing Platform

NVIDIA announced Space Computing, bringing data-center‑class AI to orbital data centers and space-edge environments via the Space‑1 Vera Rubin Module, IGX Thor, and Jetson Orin. Partners such as Aetherflux, Axiom Space, Kepler, Planet, Sophia Space and Starcloud will power next‑gen space missions with on‑board AI, real‑time analytics and autonomous operations. The lineup also includes ground‑processing acceleration with RTX PRO 6000 Blackwell Server Edition, while Space‑1 Rubin promises up to ~25x more AI compute for space inferencing versus the H100; IGX Thor and Jetson Orin are available now, with Space‑1 Rubin Module coming later.

Autonomous cytopathology hits clinical-grade accuracy with real-time 3D slide tomography
technology6 months ago

Autonomous cytopathology hits clinical-grade accuracy with real-time 3D slide tomography

Researchers report a clinical-grade autonomous cytopathology platform that uses real-time 3D whole-slide tomography with edge computing and CMD-based population analysis to automatically detect and classify cervical cells, achieving single-cell AUC over 0.99 and slide-level AUCs up to 0.97 across 1,124 samples, enabling autonomous triage cytology and objective, scalable diagnostics.

Raspberry Pi's AI HAT+ 2 brings 8GB RAM to edge AI, but power and price temper expectations
technology7 months ago

Raspberry Pi's AI HAT+ 2 brings 8GB RAM to edge AI, but power and price temper expectations

Raspberry Pi's AI HAT+ 2 adds 8GB RAM and a Hailo 10H chip to run small generative AI models locally on the Pi 5 (Llama 3.2, Qwen) with training support, at a $130 price, but independent tests suggest the Pi 5 with 8GB RAM often outperforms the add-on due to its higher power budget, making the upgrade less compelling than simply using a higher-RAM Pi or the cheaper AI HAT+ variant.

Tiny hubs, big questions: is small the future of data centres?
technology7 months ago

Tiny hubs, big questions: is small the future of data centres?

Despite ongoing expansion of large data centres, experts and industry leaders are debating a future where AI runs on devices or tiny edge centres near users, reducing latency and energy use, while some propose repurposing buildings or exploring space-based data hubs; the shift faces cost, security and reliability hurdles but reflects a move toward more bespoke, local AI processing.