Muse Glimmer is designed to run on Macs and PCs equipped with a single consumer GPU. This enables the use of local agents, function calling, local coding, and LLM-as-a-judge evaluation without an internet connection.
Model ReleasesMeta Superintelligence LabsMuse Glimmer
Meta Superintelligence Labs Releases Muse Glimmer, a 30B Parameter Model for Local Agents
This article is a translation. Read the Japanese original
The company stated that this model demonstrates strong performance in agentic use cases and benchmarks compared to existing models in its size class. The release was carried out on Hugging Face, with documentation for developers provided simultaneously.
On the technical side, it employs a compact architecture and a new distillation recipe that transfers agentic reasoning from larger teacher models. Additionally, quantization has been applied to optimize inference, ensuring that latency requirements are met.
In terms of memory efficiency, the 30B model—which would require over 55GB of memory at full precision—has been compressed to approximately 4-bit precision. This allows the language model to fit within less than 20GB, enabling simultaneous execution with KV caches and image understanding encoders in hardware environments with 24GB or 32GB of memory.
Furthermore, to improve generation speed, speculative decoding featuring a lightweight drafter model based on DFlash has been implemented. This makes it possible to maintain responsiveness even during long-form inference or multi-step tool calls.
Source: Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows(HN 1209pt・639コメント)(HN Search (backfill), 2026-08-10)