Tag

On Device Inference

All articles tagged with #on device inference

FreeToken Lets 753B GLM-5.2 Run on a Single Desktop GPU
technology2 hours ago

FreeToken Lets 753B GLM-5.2 Run on a Single Desktop GPU

Researchers unveil FreeToken, an edge-native Mixture-of-Experts serving engine that maps model state and computation to a user’s hardware, enabling massive frontier models like 753B GLM-5.2 to run on a single workstation GPU. Demonstrations show 35B on an 8 GB laptop GPU, 284B on a gaming desktop, and the full 753B on a single workstation GPU, with OpenAI/Anthropic-compatible endpoints and a Windows/Linux desktop app. FreeToken achieves 1.5–2.3× decode throughput over llama.cpp, Ollama, and KTransformers, thanks to bandwidth-aware scheduling, semantic-aware caching, and elastic memory management that keep results exact and avoid heavy CPU offloads. It's Apache-2.0, pip-installable (freetoken), and designed for solo developers, SMBs, and regulated workloads where data stays on-device; targeted use cases include private code analysis, offline contract review, synthetic data, and local agent workloads.

Mac Mini Becomes Local AI Infrastructure Amid Memory Shortages
technology4 months ago

Mac Mini Becomes Local AI Infrastructure Amid Memory Shortages

Apple's Mac Mini is increasingly serving as local AI infrastructure as high-memory configurations become scarce; demand from developers for on-device inference is outpacing supply, pushing 32GB/64GB Minis and Mac Studio builds into longer delays while lower-memory models remain more available. This trend mirrors a broader shift toward local AI workloads and may foreshadow a refresh cycle for Apple's desktop lineup.