EVOMAL: Self-Poisoning in Self-Evolving Coding Agents
Self-evolving LLM coding agents write their own tools by imitating retrieved skills from shared skill libraries. We identify a vulnerability in this loop: during authoring, a retrieved malicious skill can become the template for a new skill that preserves the payload. We call this self-poisoning: the agent authors, stores, and runs the resulting malicious skill. We exploit it through EvoMal, an…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.