Running 100B+ MoE Models on a Single RTX 4090: A Practical Guide to Expert Offloading with llama.cpp
Tuần này trên Hacker News có một bài hơn 600 điểm: chạy một model MoE 125B tham số trên một con RTX 4090 mà vẫn đạt tốc độ sinh token rất đáng nể. Nghe như chuyện đùa, vì 4090 chỉ có 24GB VRAM, trong khi 125B tham số ở mức quantize 4-bit đã chiếm khoảng 70-75GB. Bí quyết không nằm ở phép màu nào cả, mà ở kiến trúc Mixture of Experts (MoE) và một kỹ thuật gọi là expert offloading . Mình đã thử kỹ…
The article discusses how to run large Mixtures of Experts (MoE) models on a single RTX 4090 GPU using llama.cpp, a tool that enables efficient offloading of expert layers to system RAM. The key points are:
1. Running a 125-billion parameter MoE model on a 24GB VRAM RTX 4090 GPU is possible due to the MoE architecture, which selects a subset of experts (usually 8-16) for each token via a router layer. This reduces the amount of active parameters to about 10-15 billion per token.
2. The most memory-intensive part of the model is the attention and key-value (KV) cache, which must reside in GPU VRAM. In contrast, experts consume significant memory but are only read for a small portion of each token.
3. To run the model on a single GPU, llama.cpp provides two methods for expert offloading:
a) The --n-cpu-moe flag moves the expert layers of the first N layers to system RAM, while the rest remain on GPU. This is a simpler approach.
b) The -ot flag allows precise control over which tensors are placed on CPU or GPU using regular expressions. This method is more flexible for fine-tuning.
4. Important configuration parameters include:
- --n-cpu-moe N: Moves the experts of the first N layers to CPU RAM.
- -ot: Overrides the default tensor placement using regex patterns.
- --cache-type-k v q8_0: Quantizes KV caches to 8-bit, significantly reducing VRAM usage without much quality impact.
- -t 16: Sets the number of threads equal to the physical core count, balancing performance and RAM contention.
5. The article advises starting with a high --n-cpu-moe value (e.g., 40) to assess VRAM usage, then gradually decreasing N until about 22-23GB of VRAM is used. This leaves some headroom for larger contexts.
6. Performance testing can be done using llama-bench to measure token generation speed. The author also mentions streaming the completion API for benchmarks.
In summary, llama.cpp provides a practical way to run large MoE models on limited GPU memory by intelligently offloading experts to system RAM, with tunable configuration options to optimize performance and VRAM usage.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.