UOPD: Uncertainty-Aware Intervention for On-Policy Distillation of Multi-Turn Agents
On-policy distillation (OPD) trains a student on its own rollouts using dense supervision from a teacher. In multi-turn environments, a mistake at a critical decision step can redirect the subsequent rollout toward poor outcomes. We use low teacher confidence on student actions to select high-uncertainty steps for correction. In a controlled ALFWorld study, a single teacher correction at a…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.