Urgent.News

What's breaking now, across thousands of outlets.

AI

The Model Changed. My Skill Didn't. The Score Still Dropped.

What agent evals taught me about moving model floors, noisy LLM judges, and treating the evaluator as part of the instrument My rule for evaluating an agent skill is deliberately asymmetric: Test the agent on the weakest model you intend to support. Choose the judge by measuring which model grades that task reliably. Those are two different jobs. For the agent under test I want the floor: if a…

The article discusses the importance of testing agent skills on the weakest model intended for support, using a representative judge to ensure consistent measurement. The author explains that evaluating an agent skill asymmetrically, testing the weakest model and choosing a reliable judge, avoids false positives when a stronger model compensates for vague instructions.

The story is set in the context of the Keep the Why project, which incorporates a repository-native skill for coding agents. The author details how a change in the LLM judge resulted in a significant drop in scores for the agent, despite the skill not changing. By recording more than just the score, such as resolved model IDs, judge-prompt hash, CLI version, and session-shape statistics, the author was able to identify the issue and avoid prematurely rewriting the skill.

The article also highlights the importance of stabilizing the skill on the weakest supported tier and re-measuring when the model within that tier changes, as newer versions may not necessarily be stronger. The author emphasizes the need to compare instruments before comparing pass counts and notes that the same alias does not necessarily equate to the same instrument.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Let the model read the invoice, not approve it: an n8n pattern for AP automation

Most "AI invoice automation" demos stop at the fun part: a model reads a PDF and spits out JSON. The hard part is what happens next. Who decides the invoice gets paid?

  • Model only reads invoice data, separate code node handles approvals
  • Workflow triggered by Gmail for each PDF invoice attachment
  • Claude converts PDF to JSON, functions normalize extracted fields

AI Tool Calling: The Model Never Runs Your Code

A customer types one sentence into a food app's support chat: Where is order 4472? Cancel it if no rider is assigned yet. The app checks the order, cancels it, and replies. Here is the strange part.

  • Customer requests order cancellation via food app chat
  • AI model processes request without writing custom code
  • AI uses "tool calling" to delegate execution to app code

Zango AI to expand presence in Portugal

UK financial services AI company Zango AI is expanding its presence in Portugal, bringing together senior leaders from some of the country’s largest financial institutions for a new report on how AI…

More from Wednesday 7 October →