{
  "id": 11596640,
  "title": "Fine-tuning a 7B model needs 112 GB. The model is only 14 GB of it.",
  "url": "https://urgent.news/2026/10/03/fine-tuning-a-7b-model-needs-112-gb-the-model-is-only-14-gb-of-it",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-03T04:00:00.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/narotra05hp/fine-tuning-a-7b-model-needs-112-gb-the-model-is-only-14-gb-of-it-ieb"
  },
  "original_language": "en",
  "account": "To fine-tune a 7B model, 112 GB of memory is required, with only 14 GB being the actual model weights. The majority of the memory goes towards optimizer state, which is 84 GB, six times the size of the weights. The model takes up 14 GB of memory when trained in fp16 precision. By using techniques like LoRA and QLoRA, you can significantly reduce the memory requirements for fine-tuning. LoRA freezes the pretrained weights and only trains small matrices, resulting in only 0.06% of the model being trainable parameters. QLoRA takes it a step further by quantizing the frozen base, reducing its memory requirement to around 3.5 GB. This allows for fine-tuning larger models, like a 65B parameter model, on a single 48GB GPU while maintaining the performance of full 16-bit fine-tuning.",
  "summary": "Ask how much memory it takes to fine-tune a 7B model and the instinct is \"the model's 14 GB in fp16, so a bit more than that\". The real figure is about 112 GB, before you've stored a single activation. The model is 14 GB of it. Once you see where the other 98 GB goes, LoRA and QLoRA stop looking like clever tricks and start looking obvious. Where the memory actually goes The accounting comes from…",
  "key_points": [
    "Fine-tuning 7B model requires 112 GB memory",
    "14 GB for actual model weights, 84 GB for optimizer state",
    "LoRA and QLoRA reduce memory needs significantly"
  ],
  "editors_take": "Techniques like LoRA and QLoRA enable fine-tuning of large models on limited memory by reducing trainable parameters and memory requirements, allowing for larger model fine-tuning on a single GPU.",
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}