Grow the Harness, Not the Context: From Strategy-Free Scaffolds to Reusable Specialist Agents
Large language model (LLM) agents often handle streams of related tasks, yet standard harnesses repeatedly ask the model to reconstruct the same control decisions inside each task's context. We study whether task feedback can instead turn recurring control into reusable executable code, while reserving LLM calls for task-specific semantic reasoning. We introduce Growing Harness, a failure-guided…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.