ParanoiaEval: Benchmarking Unnecessary Defensive Work in Agentic Coding
As coding agents increasingly undertake real-world work autonomously, judging whether their risk treatments are warranted has become important. Existing work evaluates related agent behaviors from separate perspectives, but lacks a systematic framework for unifying these behaviors. To bridge this gap, we introduce ParanoiaEval, the first benchmark for unified evaluation of risk-treatment…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.