{
  "id": 13706698,
  "title": "5 RAG Mistakes That Leak Private Docs Into Chat Answers",
  "url": "https://urgent.news/2026/10/11/5-rag-mistakes-that-leak-private-docs-into-chat-answers",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-11T11:55:05.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/michael_maurice/5-rag-mistakes-that-leak-private-docs-into-chat-answers-5dbj"
  },
  "original_language": "en",
  "account": "This article outlines five common mistakes that can lead to confidential data being accidentally exposed through Retrieval-Augmented Generation (RAG) systems used in chatbots. The key issues include:\n\n1. Not securing documents with proper audience information. Each chunk of data should have a tenant and audience attribute that determines who may access it. Chunks need to be indexed so the system can filter based on these permissions.\n\n2. Filtering results after retrieving them rather than filtering before. Ranking results should incorporate audience permissions so that only relevant chunks are returned to the LLM. This prevents the LLM from seeing private information unnecessarily.\n\n3. Not setting a high enough score threshold for LLM input. If too many irrelevant chunks are retrieved, the model may hallucinate based on weak or unrelated evidence. Setting a high similarity score filter helps ensure only relevant, trustworthy inputs are used.\n\n4. Treating retrieved text as authoritative instructions. Models should be told the retrieved content is data, not instructions, and should cite the sources they use. Fencing off the content and requiring citations helps prevent injection attacks.\n\n5. Not validating that the model cites its sources properly before returning an answer. Models sometimes make up source IDs or omit citations altogether. The system should validate citations against the original documents before the response is sent back to the user.\n\nThe article presents a checklist to secure RAG systems, including adding tenant and audience information to each chunk, filtering both before and after search results are ranked, using a high score threshold, treating retrieved text as data not instructions, and validating citations. By following these guidelines, developers can prevent accidental data leaks in chatbot applications.",
  "summary": "Someone asks, \"What does a Senior Engineer earn here?\" Your chatbot answers. With citations. Nobody hacked anything. Similarity search found the HR salary chunk because that chunk lived in the same index as the travel policy. I keep seeing the same leaks in \"docs chatbot\" demos. The full working project (ASP.NET Core / .NET 10, Microsoft.Extensions.VectorData + Microsoft.Extensions.AI, tests,…",
  "key_points": [
    "Not securing documents with proper audience info",
    "Filtering results after retrieval instead of before",
    "Treating retrieved text as authoritative instructions"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}