CAFE: Self-Improving Search Agents Need Co-Evolving Feedback
Outcome-supervised search agents learn when and how to retrieve evidence, but terminal rewards neither localize intermediate errors nor redirect an ongoing trajectory before those errors compound. Treating corrective feedback as a learned in-trajectory intervention couples the two roles: the agent must decide when to request and use feedback, while the critic must infer useful corrections from…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.