Tag

Glm 52

All articles tagged with #glm 52

FreeToken Lets 753B GLM-5.2 Run on a Single Desktop GPU
technology2 hours ago

FreeToken Lets 753B GLM-5.2 Run on a Single Desktop GPU

Researchers unveil FreeToken, an edge-native Mixture-of-Experts serving engine that maps model state and computation to a user’s hardware, enabling massive frontier models like 753B GLM-5.2 to run on a single workstation GPU. Demonstrations show 35B on an 8 GB laptop GPU, 284B on a gaming desktop, and the full 753B on a single workstation GPU, with OpenAI/Anthropic-compatible endpoints and a Windows/Linux desktop app. FreeToken achieves 1.5–2.3× decode throughput over llama.cpp, Ollama, and KTransformers, thanks to bandwidth-aware scheduling, semantic-aware caching, and elastic memory management that keep results exact and avoid heavy CPU offloads. It's Apache-2.0, pip-installable (freetoken), and designed for solo developers, SMBs, and regulated workloads where data stays on-device; targeted use cases include private code analysis, offline contract review, synthetic data, and local agent workloads.

Zhipu’s GLM 5.2 Opens Open-Source Lead in AI Race
technology1 month ago

Zhipu’s GLM 5.2 Opens Open-Source Lead in AI Race

Zhipu’s GLM 5.2—free to download and run on your own servers—posts near parity with Anthropic Opus 4.8 on a key agentic benchmark at about one-fifth the cost, fueling demand as token spend climbs and governments curb rivals; with Anthropic Fable and OpenAI GPT-5.6 limited, the open-source model intensifies pricing pressure on frontier labs and shifts incentives toward efficiency over raw token power.

GLM-5.2: The Open-Source Chinese AI Stirring Silicon Valley
technology2 months ago

GLM-5.2: The Open-Source Chinese AI Stirring Silicon Valley

GLM-5.2, a new open-source Chinese AI model from z.AI, is drawing Silicon Valley attention for its long-context capabilities (about 1 million tokens) and strong coding performance; early endorsements from tech leaders compare it to Claude Opus 4.8 and GPT-5.5. Like DeepSeek’s R1, GLM-5.2 is open-source, letting users run and modify it locally, which could threaten the dominance of closed models from OpenAI and Anthropic. The development underscores the US-China AI race, with China pushing cheaper, capable open-source models while the US emphasizes chips and controls; industry watchers warn the window to lock in frontier capabilities may not stay open long.