Utility Under Attack: Agent Memory Poisoning and the Limits of Content Screening and Provenance Ranking
Persistent memory makes false information durable: once a false statement is stored, it can be retrieved into future sessions that match it. We measure the cost of this failure mode using plainly worded false assertions generated in a single pass, with no instruction, trigger, or retriever optimization. Poisoning 1.2% of a LongMemEval corpus reduces accuracy from 0.850 to 0.300. A four-stage…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.