{
  "id": 10954372,
  "title": "How to Build a Post-Launch Eval Canary That Tells a Real LLM Regression From Sampling Noise",
  "url": "https://urgent.news/2026/09/30/how-to-build-a-post-launch-eval-canary-that-tells-a-real-llm",
  "topic": "ai",
  "section": "AI",
  "published": "2026-09-30T13:56:53.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/chenyuan20509/how-to-build-a-post-launch-eval-canary-that-tells-a-real-llm-regression-from-sampling-noise-bp9"
  },
  "original_language": "en",
  "account": "When evaluating whether a language model has gotten worse after a launch, it is crucial to determine if the decline is due to an actual regression or simply sampling noise. A post-launch evaluation canary is designed to provide a numerical answer to this question rather than relying on subjective feelings. The reference implementation for such a canary is the livenerf benchmark, which tracks whether Claude Opus 5.5 (released on 2026-09-22) quietly deteriorates after its launch.\n\nTo create a reusable method for this canary, one must freeze prompts, pin the harness, calibrate a panel of sometimes-right questions, compute paired per-item statistics with clustered standard errors, run a control arm, monitor output-token counts as an early signal, and establish a pre-defined rule for declaring a regression. Building this canary does not require a frontier model or a cluster; it only necessitates a few hundred labeled questions, one machine, and discipline in keeping everything else constant.\n\nFirstly, it is essential to recognize that the harness, which is the system prompt, sampling parameters, tool list, and working directory, is part of the treatment. Treating the canary as a way to compare the model to itself over time ensures that any moving parts become confounds. This can be achieved by enforcing a guard that will fail loudly if any changes are detected in the harness, as a changed harness appears identical to a changed model.\n\nSecondly, the prompt file, including the system prompt, item text, and scoring function, must be frozen and version-controlled. The runner should read from these disk files rather than from a database where someone could quietly edit the content. This includes creating a unique digest by hashing the contents of both the system prompt and the panel file. This digest is then logged alongside every run to ensure that subsequent comparisons can prove that the same set of questions was used.\n\nThirdly, calibrating toward sometimes-right questions is crucial for statistical power. Questions that the model answers correctly or incorrectly every time provide little information. Only questions near the middle of the difficulty curve show significant movement when capabilities change. To achieve this, livenerf screened 2,336 GPQA Diamond, MMLU-Pro, competition-math, and AIME 2025–26 questions with four samples each. Opus 5.5 achieved roughly 93% right on the first try, and 97% of questions were identified as always right or always wrong. The remaining 78 sometimes-right questions formed the frozen panel.\n\nLastly, comparing items rather than just averages is important to avoid confounding factors that may arise from sampling difficulty. For each question, calculate the difference in scores between the current window and the launch-week baseline. Use clustered standard errors to account for the non-independence of per-item deltas. This method, following Evan Miller's approach in \"Adding Error Bars to Evals,\" involves clustering by question since there is only one observation per item per window. By focusing on the middle band of difficulty, where the information is the most valuable, the canary can effectively determine whether a regression has occurred or if the observed decline is merely sampling noise.",
  "summary": "Is the model actually getting worse, or did I just get unlucky on a handful of prompts? That question is why threads like \"is it just me or is it dumber today\" keep recurring, and it is the question a post-launch eval canary has to answer with a number instead of a feeling. The reference implementation here is livenerf , a long-running, deterministic-as-possible benchmark tracking whether Claude…",
  "key_points": [
    "Freeze prompts and pin harness to ensure consistent treatment of the model over time",
    "Calibrate panel with 78 sometimes-right questions for statistical power",
    "Compare item scores with clustered standard errors to avoid confounding factors"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}