{
  "id": 12743695,
  "title": "My Eval Said RAG Made Things Up. My Eval Was Wrong.",
  "url": "https://urgent.news/2026/10/08/my-eval-said-rag-made-things-up-my-eval-was-wrong",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-08T00:11:55.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/topunix/my-eval-said-rag-made-things-up-my-eval-was-wrong-47e"
  },
  "original_language": "en",
  "account": "A Django middleware developer created an evaluation system to test whether using an LLM with relevant source code (RAG-on) would produce more accurate explanations than using the LLM without source code (RAG-off). The system ran against a small Django blog app with 15 fixtures, each representing a different type of unhandled exception.\n\nThe evaluator, a separate model (Claude Sonnet), compared the explanations from GPT-4o-mini in both RAG modes and scored them based on whether they correctly identified the cause of the error, pointed to the fix location, proposed a working fix, and explained it clearly for a learner. The evaluator did not check for fabrication as an explicit criterion initially, but it appeared to penalize explanations that made statements not supported by the provided information.\n\nThe first iteration of the evaluation revealed that RAG-on tended to fabricate details, often inventing function parameters that were present in the source code but not mentioned in the traceback. The evaluator flagged these fabricated details as errors, even though RAG-on was supposed to use the source code to produce more accurate explanations.\n\nTo address this issue, the evaluator's criteria were rephrased to focus on whether the explanation contradicted or was unverified by the source code, rather than explicitly checking for fabrication. However, even with this change, the core problem remained: a judge with limited context could incorrectly penalize RAG-on for the very thing its purpose was to mitigate.\n\nFurther iterations involved providing the judge with more relevant source code, such as the failing function itself and any related methods. This allowed the judge to verify claims made by RAG-on more effectively. As a result, RAG-on's performance improved, particularly for fixtures where the cause of the error was in the app's code rather than in the traceback itself.\n\nOne surprising finding was that RAG-off tended to invent plausible function signatures when the traceback did not clearly specify them, while RAG-on's main issue was inventing details that were contradicted by the provided source code. Another issue discovered was that the evaluation system was truncating long tracebacks, which could prevent RAG-off from accurately identifying the location of the error in the code.\n\nIn conclusion, the evaluation revealed that while RAG-on can produce more accurate explanations when given relevant source code, the evaluation system itself needed significant improvements to accurately assess its performance. The experience highlighted the importance of providing evaluators with sufficient context to make fair judgments and the potential pitfalls of truncating tracebacks, which can obscure important information about where errors occur in the code.",
  "summary": "I maintain django-explain-errors , a Django middleware that catches unhandled exceptions in development and asks an LLM to explain them. It works in two modes: RAG-off (default): the model sees only the Django traceback. RAG-on: the model also sees relevant source code from your project, retrieved from a local sqlite-vec index by vector similarity. The whole argument for RAG-on is grounding. If…",
  "key_points": [
    "RAG-on system fabricated details in explanations",
    "Evaluation criteria revised to focus on contradiction with source code",
    "Improving source code access significantly boosted RAG-on performance"
  ],
  "editors_take": "The evaluation system's initial shortcomings highlight the challenge of fairly assessing AI models that incorporate external information, and underscore the need for evaluators to have sufficient context to make accurate judgments.",
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}