{
  "id": 30483,
  "title": "How dotdotgod Query Finds Relevant Documents from a Natural-Language Question",
  "url": "https://urgent.news/2026/08/02/how-dotdotgod-query-finds-relevant-documents-from-a-natural-language",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-08-02T05:29:43.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/dotdotgod/how-dotdotgod-query-finds-relevant-documents-from-a-natural-language-question-43do"
  },
  "original_language": "en",
  "account": "The dotdotgod query system quickly finds relevant documents from a natural-language question by leveraging the semantic relationships between the question and document passages. When the agent does not know the exact path to the relevant information, it accepts a free-form question and searches for semantically related Markdown passages within the docs/ directory. The search corpus is defined by the load.documentationSummary.exclude policy, which excludes certain bodies such as docs/plan/ and docs/archive/ to maintain the roles of current shared documentation and local working records.\n\nThe Markdown documents are split along their heading hierarchy, with each body fragment limited to 1,600 characters and attached with path and heading information. The query system then runs the multilingual E5 model locally using the @huggingface/transformers library, converting both the question and passage into normalized 384-dimensional float32 vectors. The model compares the semantic distance between the question and every document passage, adding a small bonus for matching words in titles and paths. This semantic proximity remains the primary signal, while explicit filename and heading matches also contribute to the final score.\n\nThe query system deduplicates results by Markdown path, ensuring that only the highest-scoring passage from each document is returned. Derived vector data is stored in a repository-specific cache, which is excluded from Git and rebuilt if damaged or incompatible with the current schema, model, or dimensions. The default human-readable output is concise, providing a chunk ID, repository-relative Markdown path, heading hierarchy, bounded body excerpt, and the original semantic-similarity score. This allows the agent to quickly review a bounded set of candidate sources for further reading.",
  "summary": "The dotdotgod query system enables the retrieval of relevant documents from a natural-language question by searching locally for semantically close document passages. This system accepts free-form questions and finds related Markdown passages, allowing the agent to be routed to the maintained sources worth reading. The search corpus is limited to Markdown documents under the docs/ directory, excluding active plans and historical records, while preserving the different roles of current shared documentation and local working records. The E5 model is used to convert query and passage inputs into normalized 384-dimensional float32 vectors for comparison.",
  "key_points": [
    "dotdotgod query system finds relevant documents via natural-language question",
    "Markdown passages searched semantically, excluding certain bodies",
    "E5 model compares semantic distance, adds title/path bonus"
  ],
  "editors_take": null,
  "illustration": "https://urgent.news/ill/30483.png",
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}