Urgent.News

What's breaking now, across thousands of outlets.

AI

My Eval Said RAG Made Things Up. My Eval Was Wrong.

I maintain django-explain-errors , a Django middleware that catches unhandled exceptions in development and asks an LLM to explain them. It works in two modes: RAG-off (default): the model sees only the Django traceback. RAG-on: the model also sees relevant source code from your project, retrieved from a local sqlite-vec index by vector similarity. The whole argument for RAG-on is grounding. If…

A Django middleware developer created an evaluation system to test whether using an LLM with relevant source code (RAG-on) would produce more accurate explanations than using the LLM without source code (RAG-off). The system ran against a small Django blog app with 15 fixtures, each representing a different type of unhandled exception.

The evaluator, a separate model (Claude Sonnet), compared the explanations from GPT-4o-mini in both RAG modes and scored them based on whether they correctly identified the cause of the error, pointed to the fix location, proposed a working fix, and explained it clearly for a learner. The evaluator did not check for fabrication as an explicit criterion initially, but it appeared to penalize explanations that made statements not supported by the provided information.

The first iteration of the evaluation revealed that RAG-on tended to fabricate details, often inventing function parameters that were present in the source code but not mentioned in the traceback. The evaluator flagged these fabricated details as errors, even though RAG-on was supposed to use the source code to produce more accurate explanations.

To address this issue, the evaluator's criteria were rephrased to focus on whether the explanation contradicted or was unverified by the source code, rather than explicitly checking for fabrication. However, even with this change, the core problem remained: a judge with limited context could incorrectly penalize RAG-on for the very thing its purpose was to mitigate.

Further iterations involved providing the judge with more relevant source code, such as the failing function itself and any related methods. This allowed the judge to verify claims made by RAG-on more effectively. As a result, RAG-on's performance improved, particularly for fixtures where the cause of the error was in the app's code rather than in the traceback itself.

One surprising finding was that RAG-off tended to invent plausible function signatures when the traceback did not clearly specify them, while RAG-on's main issue was inventing details that were contradicted by the provided source code. Another issue discovered was that the evaluation system was truncating long tracebacks, which could prevent RAG-off from accurately identifying the location of the error in the code.

In conclusion, the evaluation revealed that while RAG-on can produce more accurate explanations when given relevant source code, the evaluation system itself needed significant improvements to accurately assess its performance. The experience highlighted the importance of providing evaluators with sufficient context to make fair judgments and the potential pitfalls of truncating tracebacks, which can obscure important information about where errors occur in the code.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Financial ROI in AI Architecture Build vs Buy Framework for Financial Operations

Evaluating generative AI as infrastructure requires a structured build-versus-buy framework to calculate accurate Total Cost of Ownership (TCO) and maximize financial ROI.

  • Structured framework calculates Total Cost of Ownership (TCO) for AI infrastructure.
  • Build vs buy decision hinges on data sensitivity, latency, and CapEx vs OpEx considerations.

More from Thursday 8 October →