Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer
We investigate whether natural-language documentation helps coding agents resolve software issues, and we build the tools to construct and evaluate it. We introduce a roundtrip benchmark that scores code descriptions by whether code regenerated from them passes the original tests, and show that completeness, not length, drives a description's fidelity. Using the benchmark as an optimization…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.
