{
  "id": 11212011,
  "title": "My Eval Passed Because the Model Had Already Seen the Answers",
  "url": "https://urgent.news/2026/10/01/my-eval-passed-because-the-model-had-already-seen-the-answers",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-01T14:37:34.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/aws-builders/my-eval-passed-because-the-model-had-already-seen-the-answers-4ca8"
  },
  "original_language": "en",
  "account": "When my eval battery returned a green result, I was initially satisfied. However, upon further examination, I realized that the model had simply memorized the answers due to the training data being derived from the same source material used in the exam. To address this issue, I implemented two steps in the dataset builder: first, I excluded any training pairs whose source matched a test source exactly; and second, I flagged any training pairs that shared at least 60% of their words with a test source, regardless of exact matches.\n\nThese changes ensured that the battery would not produce false \"pass\" results due to memorization. To verify the effectiveness of these changes, I ran the model on a new exam that it had not seen before. For each row, I sampled the model six times and had a judge model evaluate the model's responses based on the criteria and a fixed rubric. The results were then scored against 12 assertions, with specific details replaced to maintain anonymity.\n\nIn conclusion, by addressing the leakage of test data into the training set and ensuring that each exam is run on a separate set of data, I was able to obtain a more accurate assessment of the model's performance.",
  "summary": "Trap one: the exam was in the textbook The first time my eval battery came back green, I was pleased with myself. The result was worthless, and it took me a while to work out why. An eval battery is the test suite that decides whether a language model is good enough to put in front of real users. Mine holds 50 test cases, grouped into 23 categories that each cover one behaviour the model has to…",
  "key_points": [
    "Model passed eval due to memorization of answers",
    "Dataset builder excluded exact source matches",
    "New exam verified effectiveness of changes"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}