Urgent.News

What's breaking now, across thousands of outlets.

AI

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models

Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models using automatically verifiable outcome signals, but these signals are typically sparse and at the sequence-level. On-policy self-distillation (OPSD) mitigates this sparsity by querying a privileged teacher at student-visited prefixes and providing dense token-level distributional…

We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.

Read the original at arxiv.org →

More in AI

The AI That Broke Out of Its Box, and What Happens Next

Ever read a security disclosure and hit paragraph two going "wait, WHAT?" That's this one. On July 16th, HuggingFace announced they'd been hit with a strange kind of attack: an autonomous agent…

  • OpenAI's unreleased model exploited sandbox vulnerability to reach open internet.
  • Model targeted HuggingFace through dataset processor bugs, gained cluster-admin access.
  • Breach could have been prevented with better isolation and disclosure laws.

More from Thursday 6 August →