Urgent.News

What's breaking now, across thousands of outlets.

More in AI

Running LLMs Without a GPU: What's Actually Possible

The honest numbers on CPU, NPU, and iGPU inference — what runs, how fast, and when to stop pretending. Last year a friend who runs a small consulting firm in Dubai asked me what GPU he needed to "run…

  • Running LLMs on CPU possible, performance varies by hardware
  • Token generation rates: 10-12 tokens/sec on 2021 laptop, 25-35 on M2/M3
  • Desktop CPU with 16 cores can run 8B model at 15-20 tokens/sec

More from Monday 31 August →