MoE models changed what "fits on my GPU" means, and most hardware advice hasn't caught up
My 10-year-old GPU runs 26B LLMs now, and MoE models changed everything about LLM-hosting advice
We haven't written up this one. XDA Developers has the full story — the link below goes straight to it.