Urgent.News

What's breaking now, across thousands of outlets.

AI

Building a Full RAG Pipeline in Spring Boot — From Ingest to Grounded Answer

Previous articles covered the individual pieces of RAG. You now know what embeddings are, how to chunk documents, and how Pinecone stores and searches vectors. This article connects those pieces into a working pipeline in Spring Boot — and shows exactly what happens at each step when a real question comes in. Two flows, one shared store Every RAG system has two separate flows: Ingest — runs once…

RAG systems consist of two separate flows: Ingest and Query. The Ingest flow runs once per document, splitting raw text into chunks, converting chunks into vectors, and storing them in a Pinecone vector store. The Query flow runs on every user message, converting the question into a vector, finding the most relevant chunks, injecting them into the prompt, and sending the enriched prompt to the model. The two flows share only the vector store.

The IngestService class handles the Ingest flow. It has a VectorStore and TokenTextSplitter instance, with a constructor that takes the VectorStore. The ingest() method deletes existing chunks for the provided sourceId, splits the content into chunks using the TokenTextSplitter, and adds the chunks to the vector store. The split() method returns a List of Document objects, which are then added to the vector store.

The IngestController is a REST controller that handles POST requests to /api/documents. It takes an IngestRequest object with sourceId, title, and content, and calls the ingest() method of the IngestService.

The Query flow is handled by the QuestionAnswerAdvisor bean configured in the ChatConfig class. It takes a ChatModel, VectorStore, and a search request with a topK value of 5. This means that for every query, the advisor will retrieve the 5 most similar chunks from the vector store. The system prompt instruction in the ChatClient bean ensures that the model only uses the provided context for answering and does not rely on its own knowledge.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

More from Sunday 11 October →