Urgent.News

What's breaking now, across thousands of outlets.

AI

Designing a Production RAG Retrieval System for 500 Million Vectors

Designing a Production RAG Retrieval System for 500 Million Vectors Assume there is a fictional company called Atlas. Atlas started as a fairly typical enterprise RAG product. Companies uploaded internal documents, Atlas converted them into embeddings, stored them in a vector index, and used retrieval to provide relevant context to an LLM. With around 10 million vectors, the architecture was…

Atlas, a RAG product, started simple but has grown to handle 10 million vectors. As customer data and search traffic increased, Atlas needed a more robust retrieval system. The goal was to design a system that could handle 500 million vectors without compromising performance or reliability.

The current architecture involves chunking documents, generating embeddings, storing vectors with metadata, and using these vectors for retrieval. This setup works for a smaller corpus but faces challenges when scaling up.

One issue arises when trying to maintain a single search unit for the entire dataset. A 768-dimensional float32 vector takes about 3 KB, making the raw vector data for 500 million vectors roughly 1.5 TB. This would exceed the capacity of a single machine, making it difficult to justify keeping everything in one place.

To address this, Atlas implemented a shard system, where the corpus is split into smaller parts, each with its own local search index. This allows for more efficient use of resources and easier scaling as the dataset grows. Additionally, the hottest search state can remain in memory while larger vector data and index structures can be moved to cheaper storage like SSDs.

However, this shard-based approach also creates a new challenge. As more documents are uploaded throughout the day, the write path needs to modify the optimized index in real-time. This can interfere with the read path, which involves finding relevant vectors, fetching corresponding chunks, and providing them to the LLM.

In summary, as Atlas grows to handle 500 million vectors, it faces capacity limitations and the need to separate read and write paths. The solution involves splitting the corpus into shards, each with its own local search index, and optimizing storage for different types of data. This redesign allows Atlas to maintain performance and reliability at a much larger scale.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

I built a Study Buddy for my brother who always forgets what he studied 10 mins ago ..... like literally 😭

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend What I Built I built Study Buddy, a personalized AI study assistant for my brother who usually understand stuff easily…

  • Study Buddy is an AI study assistant for retaining information
  • Created for brother who forgets studied material 10 mins later
  • Uses Gemma 4 model for personalized explanations and quizzes

Recall: A Local-First AI Memory for Everything You’ve Saved

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend What I Built I built Recall , a local-first personal memory system for a friend who has a habit of saving everything…

  • Recall is a local-first AI memory system for organizing personal files.
  • Users search memories using natural language instead of keywords.
  • The system processes files locally without cloud accounts or AI during queries.

Preserving Grandma's Recipes & Stories: Building a 100% Private, Source-Grounded Family Cookbook with Open-Source AI

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend What I Built Every family has a treasure trove of secret recipes, culinary tips, and cherished memories.

  • Grandma's Kitchen preserves family recipes using open-source AI
  • Audio recordings transcribed locally with Whisper model
  • Printable family cookbook feature creates physical heirloom cookbooks

The $250 AI Stack Is the Easy Part

These articles come from lessons learned while building Eterna Clarity and the operating system I use to run it. I can now run the recurring software stack behind Eterna for about CAD $250 a month.

What is AI-powered automation?

AI-powered automation is software that reads business documents the way a person does, then acts on them. It can read an invoice in any layout, match it to the purchase order, post it to your…

More from Sunday 4 October →