Repair or Resample? Rethinking Failure Debugging in LLM Multi-Agent Systems
As large language model (LLM)-based multi-agent systems (MASs) are increasingly applied to long-horizon complex tasks, their reliability has emerged as the core bottleneck hindering their real-world deployment. Existing MAS debugging and repair methods typically rely on rerunning and resampling the entire execution trajectory. However, a fundamental question remains to be answered: do these…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.