{
  "id": 10890045,
  "title": "VRAM for local LLMs: why memory bandwidth sets your tokens per second",
  "url": "https://urgent.news/2026/09/30/vram-for-local-llms-why-memory-bandwidth-sets-your-tokens-per-second",
  "topic": "ai",
  "section": "AI",
  "published": "2026-09-30T07:46:44.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/axrisi/vram-for-local-llms-why-memory-bandwidth-sets-your-tokens-per-second-h4h"
  },
  "original_language": "en",
  "account": "Memory bandwidth, not VRAM size, determines the speed of local LLMs. This is because generating each token requires reading all model weights from memory, making memory bandwidth the crucial factor in determining token generation speed. The ratio between memory bandwidth and model size in memory determines the maximum number of tokens per second. For instance, a 10 GB model on a 936 GB/s RTX 3090 can generate around 90 tokens in a second. The CPU's role in feeding the GPU is minimal, with a 6-core Ryzen 5 being sufficient. If a model's size exceeds VRAM, it spills into DDR5 memory, significantly reducing speed. For a 32B model, weights and KV cache combined require approximately 25 GB of VRAM. Budget tiers for local LLMs include a starter tier with an RTX 4060 Ti 16GB, a sweet spot tier with a RTX 4070 Ti Super or used RTX 3090, and an advanced tier with dual RTX 3090s or a Mac Studio.",
  "summary": "How much VRAM for an LLM is the wrong first question. The better one is how fast that VRAM is, because a local model generating text reads its entire set of weights from memory for every single token. That makes memory bandwidth, in gigabytes per second, the number that decides whether your coding agent types or crawls. This is the bandwidth-first companion to my Local AI Hardware Guide (2026)…",
  "key_points": [
    "Memory bandwidth, not VRAM size, determines LLM speed",
    "Token generation requires reading all model weights from memory",
    "RTX 3090's 936 GB/s bandwidth generates ~90 tokens/second"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}