Forgetting Only What Matters: Layer-Selective Unlearning toward Robust LLMs
Large Language Models (LLMs) can memorize and reproduce sensitive, copyrighted, or otherwise undesirable training content, creating privacy, safety, and regulatory concerns. Machine unlearning offers a practical alternative to full retraining, but many existing methods apply broad or fixed parameter updates that can degrade utility and remain brittle under deployment changes such as post-training…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.