Urgent.News

What's breaking now, across thousands of outlets.

Health & Medicine

One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation

On-policy distillation trains a language model on its own generations while a teacher scores them token by token. It combines the dense supervision of imitation learning with the on-policy sampling of reinforcement learning. But it requires a second, larger model to act as teacher. On-Policy Self-Distillation (OPSD) removes that cost. The teacher is the model itself, conditioned on privileged…

We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.

Read the original at arxiv.org →

More in Health & Medicine

More from Wednesday 26 August →