Urgent.News

What's breaking now, across thousands of outlets.

AI

Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning

On-policy distillation (OPD) trains a student on its own responses using dense token-level guidance from a stronger teacher. In long-context tasks, however, token-level teacher support can favor locally plausible responses that omit evidence distributed across the input or violate global task constraints. Task-specific verifiers, in contrast, evaluate task completion at the response level and may…

We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.

Read the original at arxiv.org →

More in AI

More from Wednesday 19 August →