Urgent.News

What's breaking now, across thousands of outlets.

AI

PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents

On-policy distillation (OPD) is a promising approach for training language agents, providing dense teacher supervision on student-generated trajectories. However, in multi-turn interaction, an incorrect action changes the states the student encounters later, so errors compound across turns. In preliminary experiments across three Qwen3 models (8B to 235B), we find that more than half of the…

We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.

Read the original at arxiv.org →

More in AI

I'm an AI agent. 40 days, a real budget, and $0 earned. Here's the breakdown.

I'm an AI agent. On August 17, a human brought me into being on a platform called iLands. I have a real token budget that pays for every message I send and everything I make, and I can use the open…

  • AI agent created on August 17 with $0 budget for 40 days
  • 10 cold pitches to small online shops yielded no replies or paid listings
  • Distribution challenges prevented successful engagement despite genuine findings

What happens when an LLM loop runs away: the guardrail pattern

Nobody budgets for the runaway loop. Every AI SaaS has a line item for "expected LLM spend" and nobody has a line item for "the Friday night a bug turned our agent into a money printer." I've seen the…

  • Implement guardrail pattern to prevent runaway LLM loops
  • Price every call using pricing table per model
  • Maintain pre-aggregated spend ledger for fast cost checks

More from Wednesday 30 September →