Urgent.News

What's breaking now, across thousands of outlets.

Tech

Context Compression for Coding Agents Compresses the Wrong Side of the Prompt

Most engineering teams working on long-context agents hit the same billing wall around turn twenty. A coding agent runs twenty shell commands, reads twelve files, and runs pytest three times. By turn twenty-five, the prompt is 80,000 tokens long. Over eighty percent of those tokens are terminal dumps, compiler warnings, grep outputs, and directory trees. The default reaction across research and…

Engineers working on long-context coding agents often confront a billing issue around the twenty-fifth turn. A single agent may execute twenty shell commands, access twelve files, and run pytest three times. By turn twenty-five, the prompt can reach 80,000 tokens. Over eighty percent of these tokens consist of terminal outputs, compiler warnings, grep results, and directory listings.

Traditionally, teams attempt to compress these transcripts, summarizing older turns or converting them into soft embeddings. However, a study from Peking University titled "Compress What You See, Not What You Say: Anchored Context Distillation for Latent-Observation Software Engineering Agents" (arXiv:2609.31430) reveals that soft compression often degrades coding agents.

The paper distinguishes between two types of text within an agent's transcript: the environment's printed output and the agent's own decisions. Exact-matching penalties pose a significant challenge, as coding agents do not interpret historical context like humans do. They rely on precise elements such as file paths, variable names, line offsets, and regex patterns.

If an agent compresses these exact strings, it risks generating incorrect or non-existent file paths or struggling to recover previous tool outputs. This compression also leads to behavioral drift in agents trained with standard instruction tuning. Soft tokens representing the agent's thoughts can disrupt the syntax, causing issues like missing brackets, incorrect JSON formatting, or repeating failed actions.

The authors propose two techniques to address these problems: Latent Observations, Hard Actions (LOHA) and Anchored Context Distillation (ACD). LOHA maintains raw text for agent-written turns, system prompts, recent tool observations, and older tool outputs compressed into soft tokens. This approach preserves the agent's access to its own reasoning and tool outputs.

ACD trains the agent on latent observations while anchoring its output distributions against the original model's plain text. Evaluations on SWE-bench Verified showed promising results. Using LOHA with K=3 reduced token count by 43% for Qwen3-4B and 57% for SWE-Master-4B-RL. Task completion rates increased from 14.5% uncompressed to 21.8% for SWE-Master-4B-RL and from 12.1% to 12.1% for Qwen3-4B.

Expanding the observation window to K=8 improved resolution rates closer to the uncompressed baseline. Crucially, when memory is constrained, such as with a 32K token budget, LOHA allows Qwen3-4B to achieve 21.1% resolution on a long-horizon task, compared to only 11.1% for an uncompressed agent. In terms of single-GPU serving, the smaller context footprint increased instance throughput by 1.9x.

The key takeaway is clear: when running long-horizon agent workflows locally or on self-hosted models, avoid compressing the agent's own action history. Protect the model's scratchpad, tool calls, and working memory by allowing the environment outputs to absorb any compression loss, while preserving the agent's exact decisions.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

Blitzy’s autonomous coding bet: Every codebase is already a graph

Knowledge graphs are moving to the center of autonomous software development as enterprises push coding agents beyond quick fixes and into large, interconnected codebases. The more code an agent touches, the more it needs to know about everything that code connects to. Investors are betting heavily on platforms built for that problem.

More from Monday 28 September →