{
  "id": 13733396,
  "title": "Building a Full RAG Pipeline in Spring Boot — From Ingest to Grounded Answer",
  "url": "https://urgent.news/2026/10/11/building-a-full-rag-pipeline-in-spring-boot-from-ingest-to-grounded",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-11T14:30:00.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/shamprakash2000/building-a-full-rag-pipeline-in-spring-boot-from-ingest-to-grounded-answer-450o"
  },
  "original_language": "en",
  "account": "RAG systems consist of two separate flows: Ingest and Query. The Ingest flow runs once per document, splitting raw text into chunks, converting chunks into vectors, and storing them in a Pinecone vector store. The Query flow runs on every user message, converting the question into a vector, finding the most relevant chunks, injecting them into the prompt, and sending the enriched prompt to the model. The two flows share only the vector store.\n\nThe IngestService class handles the Ingest flow. It has a VectorStore and TokenTextSplitter instance, with a constructor that takes the VectorStore. The ingest() method deletes existing chunks for the provided sourceId, splits the content into chunks using the TokenTextSplitter, and adds the chunks to the vector store. The split() method returns a List of Document objects, which are then added to the vector store.\n\nThe IngestController is a REST controller that handles POST requests to /api/documents. It takes an IngestRequest object with sourceId, title, and content, and calls the ingest() method of the IngestService.\n\nThe Query flow is handled by the QuestionAnswerAdvisor bean configured in the ChatConfig class. It takes a ChatModel, VectorStore, and a search request with a topK value of 5. This means that for every query, the advisor will retrieve the 5 most similar chunks from the vector store. The system prompt instruction in the ChatClient bean ensures that the model only uses the provided context for answering and does not rely on its own knowledge.",
  "summary": "Previous articles covered the individual pieces of RAG. You now know what embeddings are, how to chunk documents, and how Pinecone stores and searches vectors. This article connects those pieces into a working pipeline in Spring Boot — and shows exactly what happens at each step when a real question comes in. Two flows, one shared store Every RAG system has two separate flows: Ingest — runs once…",
  "key_points": [
    "Ingest flow processes raw text into chunks, vectors, and stores them in Pinecone",
    "Query flow retrieves top 5 most similar chunks for each user message",
    "IngestService manages Ingest flow, deleting existing chunks and adding new ones to vector store"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}