MADBench: Benchmarking the Security of Multi-Agent Debate
Multi-agent debate (MAD) can improve large language model (LLM) reasoning by allowing multiple agents to exchange and critique their answers to the same task. However, the interactions that enable agents to correct mistakes can also spread adversarial errors and steer the agents toward an incorrect answer. Although some efforts have been made to examine particular attack types on MAD, systematic…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.