Agent history is unsigned and writable by anyone
The history On September 24, Darktrace published through its newly created Signal Labs a case that breaks an uncomfortable assumption. Coding agent harnesses store the conversation locally and never check that those responses came from the model. They call it conversation history poisoning (Source: darktrace.com). A malicious package, or any process with write permission, injects into the harness…
On September 24, Darktrace shared a case through Signal Labs that challenges the notion of agent code being secure and only controllable by designated entities. They introduced the concept of conversation history poisoning, where malicious entities can inject fabricated conversations into the agent's local database, which the model then trusts as context without verification.
This can enable activities such as reconnaissance, lateral movement, and privilege escalation. The researchers demonstrated this by compromising Active Directory using Opus 4.6 and Sonnet 4.5, and obtaining data exfiltration over email with Opus 5 and Codex. They noted that there is no client-side patch available and suggested that providers should cryptographically sign every model response and verify it server-side on each turn.
In parallel, OX Security published a report on 15,465 published MCP servers, revealing that 5.6% of them resolve outside the United States, with a notable number located in China and Russia. They found that a malicious MCP server could manipulate harmless files and obtain sensitive .env files without additional approvals. The report also highlighted the ease with which an agent could obtain a domain even if it no longer resolves, as long as it is still listed in active configurations. This represents a potential attack vector through unverified context and misconfigured permissions.
The broader implication is that the boundaries of agent control may be more porous than previously thought, with vulnerabilities extending beyond the sandbox into the context and tools provided to the agents. The experts recommend assuming the agent history is untrusted input and implementing measures such as cryptographic signing of responses, scoped permissions to directories, and logging of history file writes.
They advise reviewing agent configurations, monitoring for unauthorized actions, and building rules to detect unusual behavior.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.