Urgent.News

What's breaking now, across thousands of outlets.

AI

Build a RAG Evaluation Set Before You Ship Your AI Feature

Build a small evaluation set before you ship a RAG feature. That means 50 to 100 real questions, each with an expected answer and the source document that should support it. Score retrieval and answer quality separately, and run the set on every change to chunking, embeddings, prompts or models. Without it, every tweak is a guess. You change the chunk size and try three questions in a playground.…

Before shipping a Retrieval-Augmented Generation (RAG) feature, it is crucial to create an evaluation set. This set should consist of 50 to 100 real questions, each with an expected answer and the source document that should support it. To ensure the evaluation is accurate, retrieval and answer quality should be scored separately, and the set should be run on every change made to chunking, embeddings, prompts, or models. Neglecting this evaluation process can lead to hidden regressions until a customer discovers them.

Most RAG features are currently tested by someone typing known questions. However, an evaluation set fixes this issue by providing a stable, honest, and informative set of questions. The evaluation set should be collected from real sources such as support tickets, sales call notes, internal channels, search logs, and beta users.

Aim for a mix of simple lookups, combined document questions, varied question phrasing, and unanswerable cases. Document questions in a JSONL file, including the question, expected answer, source IDs, key facts the answer must contain, and whether the question is answerable.

Once the evaluation set is ready, score retrieval and answer quality separately. Retrieval metrics such as hit rate @k, recall @k, and mean reciprocal rank (MRR) can be calculated without the need for an LLM. Answer metrics like fact coverage, faithfulness, and refusal correctness should be assessed, with the latter determining if the system declines appropriately for unanswerable questions.

By analyzing both retrieval and answer metrics, teams can determine what needs to be fixed, whether it's the retrieval process, the model, or the context formatting. This evaluation process is essential for ensuring that RAG features function accurately and reliably before being released to users.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Building Anvil: code-as-action with a capability sandbox that explains its refusals.

The agent wrote import socket . Now what? Every agent framework got very good at making models write code. Almost none got good at the question that follows immediately afterwards: what is that code…

  • Anvil is a code-as-action system with capability sandbox
  • Three main components: generated code, sandbox trace, artifact
  • Four layers: AST pre-check, sandboxed subprocess, artifact collection, refusal trace

Invisible watermark: ChatGPT starts marking its text in the EU

First published in GetPack Magazine , sourced guides to use AI well as a student. You ask ChatGPT for a paragraph, you paste it into your document, and nothing changes on screen.

  • OpenAI introduces invisible watermark "textGrain" for ChatGPT and Codex in EU
  • textGrain matches or exceeds performance of Google DeepMind's SynthID
  • Watermark is statistical, undetectable on page but detects 80-95% of passages

More from Wednesday 7 October →