Urgent.News

What's breaking now, across thousands of outlets.

More in AI

How to Build a Post-Launch Eval Canary That Tells a Real LLM Regression From Sampling Noise

Is the model actually getting worse, or did I just get unlucky on a handful of prompts? That question is why threads like "is it just me or is it dumber today" keep recurring, and it is the question a…

  • Freeze prompts and pin harness to ensure consistent treatment of the model over time
  • Calibrate panel with 78 sometimes-right questions for statistical power
  • Compare item scores with clustered standard errors to avoid confounding factors

More from Wednesday 30 September →