Running a 180B-Parameter MoE Model on a Gaming Laptop: VIDRAFT's POCKET-Darwin-180B-GGUF
Running a 180B-Parameter MoE Model on a Gaming Laptop: VIDRAFT's POCKET-Darwin-180B-GGUF TL;DR: VIDRAFT has released POCKET-Darwin-180B-GGUF, a 4-bit quantized, GGUF-format version of their 180B Mixture-of-Experts model that runs on consumer hardware — including a gaming laptop with 8 GB VRAM and 32 GB RAM. By combining MoE sparse activation with llama.cpp-based SSD streaming, only ~3B parameters…
VIDRAFT has unveiled POCKET-Darwin-180B-GGUF, a compressed 4-bit quantized version of its 180B Mixture-of-Experts model, allowing it to run on consumer-grade hardware like a gaming laptop. By leveraging sparse MoE activation, 4-bit quantization, and llama.cpp-based SSD streaming, the model only requires ~3B active parameters at inference time, cutting hardware requirements by approximately 250 times compared to the full-precision server version.
This compressed edition retains 87.65% accuracy on MMLU-Pro, demonstrating negligible quality degradation. VIDRAFT's strategy combines the RSI (Recursive Self-Improvement) capability with POCKET deployability, enabling on-premise, air-gapped, or local inference without a cloud account or dedicated GPU server. The model is available in GGUF format on Hugging Face and ModelScope, accessible via the Hugging Face CLI.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.