Fine-tuning a 7B model needs 112 GB. The model is only 14 GB of it.
Ask how much memory it takes to fine-tune a 7B model and the instinct is "the model's 14 GB in fp16, so a bit more than that". The real figure is about 112 GB, before you've stored a single activation. The model is 14 GB of it. Once you see where the other 98 GB goes, LoRA and QLoRA stop looking like clever tricks and start looking obvious. Where the memory actually goes The accounting comes from…
To fine-tune a 7B model, 112 GB of memory is required, with only 14 GB being the actual model weights. The majority of the memory goes towards optimizer state, which is 84 GB, six times the size of the weights. The model takes up 14 GB of memory when trained in fp16 precision. By using techniques like LoRA and QLoRA, you can significantly reduce the memory requirements for fine-tuning.
LoRA freezes the pretrained weights and only trains small matrices, resulting in only 0.06% of the model being trainable parameters. QLoRA takes it a step further by quantizing the frozen base, reducing its memory requirement to around 3.5 GB. This allows for fine-tuning larger models, like a 65B parameter model, on a single 48GB GPU while maintaining the performance of full 16-bit fine-tuning.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.