{
  "id": 12596647,
  "title": "The Model Changed. My Skill Didn't. The Score Still Dropped.",
  "url": "https://urgent.news/2026/10/07/the-model-changed-my-skill-didnt-the-score-still-dropped",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-07T09:48:26.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/oliverzehentleitner/the-model-changed-my-skill-didnt-the-score-still-dropped-mlj"
  },
  "original_language": "en",
  "account": "The article discusses the importance of testing agent skills on the weakest model intended for support, using a representative judge to ensure consistent measurement. The author explains that evaluating an agent skill asymmetrically, testing the weakest model and choosing a reliable judge, avoids false positives when a stronger model compensates for vague instructions. The story is set in the context of the Keep the Why project, which incorporates a repository-native skill for coding agents. The author details how a change in the LLM judge resulted in a significant drop in scores for the agent, despite the skill not changing. By recording more than just the score, such as resolved model IDs, judge-prompt hash, CLI version, and session-shape statistics, the author was able to identify the issue and avoid prematurely rewriting the skill. The article also highlights the importance of stabilizing the skill on the weakest supported tier and re-measuring when the model within that tier changes, as newer versions may not necessarily be stronger. The author emphasizes the need to compare instruments before comparing pass counts and notes that the same alias does not necessarily equate to the same instrument.",
  "summary": "What agent evals taught me about moving model floors, noisy LLM judges, and treating the evaluator as part of the instrument My rule for evaluating an agent skill is deliberately asymmetric: Test the agent on the weakest model you intend to support. Choose the judge by measuring which model grades that task reliably. Those are two different jobs. For the agent under test I want the floor: if a…",
  "key_points": [
    "The author emphasizes testing agent skills on the weakest model for consistent measurement.",
    "A change in the LLM judge caused a significant drop in scores for the agent."
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}