The Model Changed. My Skill Didn't. The Score Still Dropped.
What agent evals taught me about moving model floors, noisy LLM judges, and treating the evaluator as part of the instrument My rule for evaluating an agent skill is deliberately asymmetric: Test the agent on the weakest model you intend to support. Choose the judge by measuring which model grades that task reliably. Those are two different jobs. For the agent under test I want the floor: if a…
The article discusses the importance of testing agent skills on the weakest model intended for support, using a representative judge to ensure consistent measurement. The author explains that evaluating an agent skill asymmetrically, testing the weakest model and choosing a reliable judge, avoids false positives when a stronger model compensates for vague instructions.
The story is set in the context of the Keep the Why project, which incorporates a repository-native skill for coding agents. The author details how a change in the LLM judge resulted in a significant drop in scores for the agent, despite the skill not changing. By recording more than just the score, such as resolved model IDs, judge-prompt hash, CLI version, and session-shape statistics, the author was able to identify the issue and avoid prematurely rewriting the skill.
The article also highlights the importance of stabilizing the skill on the weakest supported tier and re-measuring when the model within that tier changes, as newer versions may not necessarily be stronger. The author emphasizes the need to compare instruments before comparing pass counts and notes that the same alias does not necessarily equate to the same instrument.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.