Meta’s new AI model runs entirely offline, but your GPU needs to keep up
Meta released Muse Glimmer, a free 30B parameter AI model you can run entirely offline. No subscription, no data center, just a GPU with at least 24GB of VRAM.
Meta has unveiled a new AI model called Muse Glimmer, notable for its unique offline capabilities. This 30-billion-parameter model is released under an Apache 2.0 license, allowing users to download, modify, and expand upon the weights without any restrictions. Muse Glimmer operates solely on a local graphics card, eliminating the need for server farms or internet connectivity.
The model is capable of processing both text and images, delivering text-based responses. It supports over 100 languages and can maintain conversation threads exceeding 131,000 tokens. The knowledge cutoff for Muse Glimmer is set to January 4, 2026.
To accommodate the model's size, Meta employs quantization, compressing the computational requirements to under 20GB for the language portion. Two versions are available: K-Quant-Dynamic, which requires 32GB of video memory and retains 98% of the model's accuracy, and K-Quant-17GB, which fits within 24GB with a slight, 1% accuracy reduction.
Running Muse Glimmer necessitates a powerful GPU. Minimum requirements include an RTX 5090, RTX 4090, RTX 3090, or an Apple Silicon Max chip. The model also incorporates an accelerator known as DFlash, which enhances token prediction speed. On an RTX 5090, this acceleration boosts performance from 74.9 to 233.4 tokens per second, a threefold improvement over previous models.
In comparison to Google's Gemma4 and Alibaba's Qwen3.6, Muse Glimmer excels in planning and multi-step tasks. However, it lags behind in direct interaction with desktop environments.
The AI model is now available for download via Hugging Face or LM Studio, provided users possess the necessary hardware.
Written by urgent.news from Digital Trends's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.