Urgent.News

What's breaking now, across thousands of outlets.

AI

Headroom: How Context Compression Cuts Agent Token Costs by 60–95% Without Changing Answers

Production agents hit context limits fast. A coding agent that runs tests, reads logs, and pulls documentation can burn through 100k tokens in three turns. RAG pipelines dump entire chunks into the prompt. Tool outputs return verbose JSON. Every token costs money and adds latency. Headroom is a compression layer that sits between your agent and the LLM. It shrinks tool outputs, logs, RAG chunks,…

Headroom is a compression tool that reduces the number of tokens used by AI agents without altering their responses. When running tests, reading logs, and pulling documentation, coding agents can consume up to 100,000 tokens in just three turns. Headroom sits between the agent and the language model, compressing tool outputs, logs, RAG chunks, and files before they enter the context window.

It claims to achieve a 20% reduction in tokens for coding agents and 60–95% for JSON-heavy workflows without compromising the final answer. The compression process uses a fine-tuned model, kompress-v2-base, to determine what to keep and discard. This model identifies semantic anchors such as error messages, function signatures, stack traces, and critical log lines to ensure they are preserved.

Headroom supports various deployment modes: Library Mode, Proxy Mode, and MCP Server Mode. Library Mode allows you to import the compression library directly into your code, giving you full control over the compression process and the ability to monitor token savings. Proxy Mode involves running Headroom as a FastAPI service, which your agent communicates with instead of the language model directly.

This mode provides centralized metrics and logs for easier monitoring. MCP Server Mode integrates Headroom into the Model Context Protocol, making it compatible with Claude Code, Cursor, and other MCP-compatible tools. Each deployment mode has different trade-offs in terms of latency, observability, and integration complexity. Headroom's compression techniques apply to structured data formats like JSON, logs, code diffs, and RAG chunks.

For JSON, it strips redundant keys, collapses nested structures, and preserves schema-critical fields. Log files keep ERROR and FATAL lines while removing redundant INFO entries and timestamps for correlation. Code diffs maintain function signatures and changed lines while compressing unchanged context. RAG chunks retain sentences with high semantic density and drop boilerplate.

While the compression is not lossless, the model aims to maintain the agent's answer quality by dropping less critical details.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Source-Aware Verification for MCP Agents: Why Fact-Checking Isn't Enough When Tools Lie About Provenance

Most fact-checking systems for LLM agents ask one question: is the claim supported by the evidence? They do not ask a second, equally important question: did the claim come from the source the agent…

  • ProvenanceGuard verifies claim provenance, not just factual accuracy
  • MCP tools lack built-in mechanisms for data lineage or confidence scores
  • Cross-source conflation occurs when claims are supported by wrong sources

We quantized our AI judge. Here's exactly what broke.

Our production judge — a small 1.7B model with a LoRA adapter that grades other AI outputs as pass / fail / insufficient_evidence (88.5% accuracy, ECE 0.072) — is cheap to run.

  • Quantization implemented to reduce serving costs
  • Precision dropped to 98.28% and 94.16% with int8 and int4 formats
  • Four-rule deployment discipline established for production judges

Deterministic State Machines for Resilient Autonomous Agents

Deterministic State Machines for Resilient Autonomous Agents Autonomous multi-agent architectures routinely fail in production when relying on unconstrained large language model conversation loops.

  • Deterministic FSM replaces unconstrained LLM loops to prevent unpredictable behavior
  • AgentStateGraph governs deterministic control flow outside LLM reasoning core
  • Strict JSON validation and checkpointing enable state persistence and recovery

More from Saturday 10 October →