{
  "id": 13635422,
  "title": "How much RAM you actually need to run AI locally",
  "url": "https://urgent.news/2026/10/11/how-much-ram-you-actually-need-to-run-ai-locally",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-11T04:01:32.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/selfhostpilot/how-much-ram-you-actually-need-to-run-ai-locally-5hl2"
  },
  "original_language": "en",
  "account": "Higher-end AI PCs are now available at prices that prompt people to ask the incorrect first question: which model offers the best value. The crucial factor determining what you can run is not the price, the number of cores, or the TOPS figure. Instead, it is how much unified memory you can allocate to the model and how much remains after operating system overhead. Below is a sizing table that is sorely missing from most hardware reviews. Each model class has specific memory requirements: Small models (7–8B parameters) need around 5 GB at 4-bit quantization; Mid-size models (30B) require approximately 15–20 GB at 4-bit, not including context; and Large models (70B) need about 35–40 GB at 4-bit, plus context. Context encompasses elements like lengthy conversations, extensive system prompts, or repositories of code—everything that must reside in memory alongside the model weights. Neglecting context memory can lead to hitting a limitation during a session. The unified-memory dynamic on these machines is noteworthy. A 128 GB machine demonstrated roughly 110 GB of GPU-addressable memory—about 79.9 GB dedicated to the model and 30 GB shared with the CPU. This is not the total 128 GB available; it is the accurate figure to plan for. In practice, 24 GB memory becomes restrictive once the operating system claims its portion. In contrast, 64 GB and 128 GB configurations are not merely marketing fluff but are essential for running a 70B model efficiently. Before investing in an AI PC, consider three questions. First, which models do you intend to run? Choose the most demanding one honestly, then refer to the sizing table. Second, how much context will you actually utilize? A code assistant with access to a repository in context may require multiple times the weight of the model itself. Third, do you need to own the hardware, or would renting inference in the cloud suffice? Cloud rentals incur no upfront costs, allow access to models exceeding the memory capacity of these machines, and offer privacy, lower latency, offline capabilities, and better long-term economics. The price range of current AI PCs spans from $2,599 to $5,999. The difference between the cheapest and most expensive models primarily revolves around memory decisions—$24 GB at the lowest end and $128 GB at the highest. Spending more on additional cores while limiting memory to 24 GB results in a machine incapable of running the desired models. Therefore, start by selecting the desired model, calculate the required memory, and then compare prices. This order of operations typically results in a lower bill. The article originally appeared on SelfHost Pilot, where the author tests self-hosted software and AI stacks on real hardware and shares measured results. For those interested in tracking self-hosted software trends based on commit momentum rather than star counts, the author maintains a tracker with over 1,700 projects at selfhosted-tracker.",
  "summary": "Higher-end AI PCs landed this month at prices that make people ask the wrong question first: which one is worth the money? The number that actually decides what you can run is not the price, the core count, or the TOPS figure. It is how much unified memory you can give the model , and how much of it survives the operating system. Here is the sizing table I wish more hardware pages published. What…",
  "key_points": [
    "Small models (7–8B parameters) need around 5 GB at 4-bit quantization.",
    "Mid-size models (30B) require approximately 15–20 GB at 4-bit, not including context.",
    "Large models (70B) need about 35–40 GB at 4-bit, plus context."
  ],
  "editors_take": "Prioritizing memory capacity over other specs when buying an AI PC can lead to a lower bill and better performance, as the amount of unified memory determines which models can be run.",
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}