DualOPSD: Adaptive Privileged Teachers for On-Policy Self-Distillation
On-policy self-distillation (OPSD) uses a privileged copy of the student model to provide dense supervision without an external teacher. OPSD keeps this privileged teacher fixed, even though the student distribution and output style change during training. We propose DualOPSD, an asymmetric alternating framework that adapts both policies. The student first learns from the privileged teacher. The…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.