Urgent.News

What's breaking now, across thousands of outlets.

AI

The Math an Applied AI Engineer Actually Uses in an Evaluation Loop

A team compares two versions of an AI service. Average accuracy rises from 82% to 86%. The new version looks better, so releasing it seems like the obvious decision. Then an engineer separates the failures by type. The system now makes fewer mistakes on simple questions, but it is twice as likely to answer confidently without a source when the question concerns billing or contracts. The average…

In an evaluation loop for an AI service, engineers compare two versions of the system. Initially, the new version shows higher average accuracy, but further analysis reveals it is more likely to answer confidently without sources for billing or contract-related questions. This discrepancy in performance highlights the need for applied mathematics in AI engineering decisions.

To begin, engineers should construct a small evaluation loop, using mathematical formulas that directly address engineering questions. The first step is to establish a unit of evaluation. A weak test set should contain more than just query-answer pairs; it should include the query, scenario or segment, source identifiers, expected behavior, whether human review is required, and the cost of the failure. A simple Python dictionary can represent this first version of the evaluation case.

The cost associated with a failure is not a universal mathematical constant; it is a team decision based on the risk involved. For example, a mistake in a low-risk reference answer might cost 1 point, while an unsupported claim about a payment could cost 5 points. An action with irreversible consequences should be blocked by a separate policy rather than being reduced to a score.

When evaluating retrieval in a RAG system, it is essential to separate the retrieval process from the final answer. A bad answer can stem from various causes, such as the required document not being indexed, retrieval failing to return it near the top, the model receiving the correct source but ignoring it, or the source being obsolete.

Measuring only the final answer fails to distinguish between these different failure types. A useful retrieval metric is recall@k, which calculates the fraction of required sources that appear in the first k results.

To gain a more comprehensive understanding of the system's performance, engineers should also analyze failure types separately from average accuracy. For instance, in a system that classifies requests and routes them to a department, accuracy alone is insufficient when classes are imbalanced or different mistakes have different consequences.

Engineers should distinguish between various failure types, such as false positives, missed cases, unsupported answers, unnecessary refusals, and wrong escalations. By assigning a cost to each failure type, engineers can calculate the weighted risk, which provides a more nuanced view of the system's performance. The final decision on whether to release the new version, return it for further work, or restrict one scenario should be made by the process owner, based on the agreed-upon cost of each failure type.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

More from Sunday 4 October →