Urgent.News

What's breaking now, across thousands of outlets.

AI

Part 3: Knowing when your agent doesn’t know: the confidence layer

The confidence layer is a crucial aspect of running large language model (LLM) systems in production. It is the third level in a six-level maturity model, building upon the foundation of successful operation and visibility provided by levels one and two. At this level, the system autonomously acts only when its calibrated confidence is high.

Decisions that matter are graded with an independent judge, and anything uncertain is routed to a human. Confidence determines what runs automatically, with a judge acting as a signal for the automated path and a trigger for pulling decisions off it. Handing off to a human is where the confidence dial sends everything below the line.

The most important number an agent produces is not its answer but its calibrated confidence, which should be composed from independent signals, never derived from the model's own claims. This confidence signal should be calibrated, meaning that a score of 0.9 should indicate correct performance 90% of the time. To achieve this, decisions should be bucketed based on predicted confidence, and real accuracy should be measured per bucket using a reliability diagram.

If the top bucket is right 71% of the time, the automation threshold is too loose. Calibration should be closed by fitting a monotonic map from raw composed score to empirical accuracy and routing based on that. The calibration should be recomputed on a rolling window, as it drifts with the model and inputs. The Brier score or the signed per-bucket gap should be preferred for alerting over the Expected Calibration Error (ECE), as ECE can be binning-sensitive and may read 0 for a miscalibrated model.

Independence is key in the confidence layer, as weighting historical slice accuracy into compose() and then calibrating per slice double-counts. The composed and calibrated score then serves as a threshold, with the right threshold set for each slice based on the calibration data. The highest form of automation is abstention, where an agent declines to decide when it's not confident.

This is more trustworthy than always providing an answer. The three levers for tuning the system are calibration, abstain threshold, and routing decisions. Anti-patterns that can ruin this layer include using raw model confidence as the dial, employing a single global threshold, never re-checking calibration, and treating abstention as an error.

A judge serving as a second, independent model that evaluates the first's decision before acting is crucial for high-stakes decisions. The judge should be independent, skeptical, and differ from the original model in family or prompt. Disagreement between the model and judge should trigger a human review, and the rate of disagreement serves as a health signal, indicating potential attacks, bad deploys, or model regressions.

Written by urgent.news from Stack Overflow Blog's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at stackoverflow.blog →

More in AI

More from Wednesday 7 October →