{
  "id": 12068099,
  "title": "Building a RAG Pipeline for Semantic Code Search",
  "url": "https://urgent.news/2026/10/05/building-a-rag-pipeline-for-semantic-code-search",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-05T03:52:48.000Z",
  "source": {
    "name": "Lobsters",
    "slug": "lobsters",
    "url": "https://blog.jetbrains.com/ai/2026/09/building-a-rag-pipeline-for-semantic-code-search-a-developer-diary-and-field-notes/"
  },
  "original_language": "en",
  "account": "Building a RAG Pipeline for Semantic Code Search\n\nWe embarked on a mission to create a semantic code search platform using retrieval-augmented generation (RAG) technology inside JetBrains products. The result was Air Context, a production-ready solution that we refined through various iterations. In this series, we will share insights gained during the development process, starting with the foundational stages of parsing and chunking, and vectorization.\n\nParsing and chunking are essential steps in preparing raw source files for the LLM to process. While this may seem straightforward for small-scale projects, large-scale codebases pose significant challenges. Large files and extensive code can overwhelm the system, leading to irrelevant retrieval results. To mitigate this, we must divide the code into properly scoped chunks, ensuring each group contains enough context and represents a common semantic meaning. Our experience at JetBrains has given us access to intelligent parsers that can adapt to various programming languages, forming part of our internal Code Engine platform. By leveraging these parsers, we can intelligently chunk the code, striking the right balance between context and semantic relevance.",
  "summary": null,
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}