{
  "id": 10983662,
  "title": "Search Agents Waste Half Their Tokens Rediscovering Entity Links",
  "url": "https://urgent.news/2026/09/30/search-agents-waste-half-their-tokens-rediscovering-entity-links",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-09-30T16:29:02.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/reidmarlow/search-agents-waste-half-their-tokens-rediscovering-entity-links-31an"
  },
  "original_language": "en",
  "account": "In a new paper, researchers from KAIST and Microsoft reveal that search agents can waste up to half of their tokens by blindly navigating through documents while trying to find answers to questions. The issue arises because standard flat document collections provide no relational pointers between files, forcing the model to reconstruct the web of cross-document relationships at inference time. To address this problem, the researchers introduce CorpusMap, a system that structures the corpus around recurring named entities like people, projects, systems, vendors, and code modules. Offline, an extraction pipeline identifies these recurring anchors and generates Entity Pages for each one, which aggregate key facts about the entity and maintain explicit backlinks to every original document that mentions it. This creates a bipartite graph between entities and raw documents. When the agent receives a task, it navigates this graph using standard terminal commands, bypassing the need for blind keyword grepping and jumping directly to the relevant entity page. The authors demonstrate that CorpusMap significantly reduces token consumption and improves answer correctness compared to raw corpus search. On EnterpriseRAG-Bench using GPT-5.5, the token consumption drops from 206,500 tokens to 88,100 tokens (a 57% reduction) and answer correctness increases from 62.1% to 73.8%. On WixQA, token consumption drops from 337,200 tokens to 74,500 tokens (a 78% reduction) and factual accuracy rises from 67.5% to 70.7%. This approach consistently improves performance across seven different model families. The upfront indexing cost of building entity maps is a concern, but the paper shows that using a cheaper model for construction (e.g., GPT-5.6 Luna) still yields high-quality entity pages, with only a minimal impact on retrieval quality. Updating the entity graph with new documents is also efficient, as it only requires updating the touched entity pages instead of reprocessing the entire corpus.",
  "summary": "If you wire an LLM agent to a local directory of documents and give it terminal tools (grep, find, cat), you quickly notice an ugly pattern. When a question depends on evidence scattered across three separate files, the agent spends most of its trajectory wandering in circles. It greps for a keyword, pulls up five irrelevant markdown files, reads their headers, backs up, reformulates the search,…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}