Urgent.News

What's breaking now, across thousands of outlets.

AI

Fine-tuning a 7B model needs 112 GB. The model is only 14 GB of it.

Ask how much memory it takes to fine-tune a 7B model and the instinct is "the model's 14 GB in fp16, so a bit more than that". The real figure is about 112 GB, before you've stored a single activation. The model is 14 GB of it. Once you see where the other 98 GB goes, LoRA and QLoRA stop looking like clever tricks and start looking obvious. Where the memory actually goes The accounting comes from…

To fine-tune a 7B model, 112 GB of memory is required, with only 14 GB being the actual model weights. The majority of the memory goes towards optimizer state, which is 84 GB, six times the size of the weights. The model takes up 14 GB of memory when trained in fp16 precision. By using techniques like LoRA and QLoRA, you can significantly reduce the memory requirements for fine-tuning.

LoRA freezes the pretrained weights and only trains small matrices, resulting in only 0.06% of the model being trainable parameters. QLoRA takes it a step further by quantizing the frozen base, reducing its memory requirement to around 3.5 GB. This allows for fine-tuning larger models, like a 65B parameter model, on a single 48GB GPU while maintaining the performance of full 16-bit fine-tuning.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Local Models vs Cloud Models — The AI Advantage Is Moving From Access to Infrastructure

Local Models vs Cloud Models — The AI Advantage Is Moving From Access to Infrastructure Core thesis AI is no longer new. The question isn't "Are you using AI?" It's: "How much of your work has AI…

  • AI is no longer novel; focus shifts to infrastructure.
  • Domain experts and marketers/sales professionals will rise.
  • Cloud vs local models debate emerges as AI costs rise.

The More Context You Give Your AI Coding Agent, the Worse It Can Get

We keep hearing the same advice: Give the AI more context. Add the README. Add AGENTS.md . Add architecture docs. Add logs. Add previous decisions. Add the whole repository.

  • Too much context can introduce noise and stale assumptions for AI agents.
  • Conflicting instructions from different sources create confusion for AI agents.
  • Minimum sufficient context, not all available information, yields better results.

The AI Autopilot Trap: Are We Outsourcing Our Minds

Let’s be honest: we are living in the golden age of convenience. Need an email drafted? Ask AI. Need code debugged, a marketing strategy outlined, or a complex book summarized? Done in seconds.

  • AI can perform tasks like drafting emails and summarizing complex texts.
  • Daily reliance on AI weakens critical thinking and problem-solving skills.
  • Use AI as a co-pilot, retain final evaluation and decision-making as uniquely human tasks.

More from Saturday 3 October →