Urgent.News

What's breaking now, across thousands of outlets.

More in AI

Running 100B+ MoE Models on a Single RTX 4090: A Practical Guide to Expert Offloading with llama.cpp

Tuần này trên Hacker News có một bài hơn 600 điểm: chạy một model MoE 125B tham số trên một con RTX 4090 mà vẫn đạt tốc độ sinh token rất đáng nể.

  • A 125-billion parameter MoE model runs on a 24GB VRAM RTX 4090 using llama.cpp
  • llama.cpp offers two expert offloading methods: --n-cpu-moe and -ot flags
  • Quantizing KV caches to 8-bit reduces VRAM usage significantly

"Your Ai coach"

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend What I Built Code How I Built It Why Does Open Innovation Matter? Prize Categories

More from Monday 5 October →