Urgent.News

What's breaking now, across thousands of outlets.

AI

Done Is Testimony. Terminal State Is the Grade.

Two things landed in the same week and they are the same argument. Microsoft and Hugging Face published ThinkingBox , a benchmark that grades agents on the backend state and side effects they leave behind instead of the sentences they generate, and then asks whether they can do it twenty times in a row. Their opening example is an agent that makes nine clean tool calls, closes a ticket as…

Two developments in the same week share the same argument: Microsoft and Hugging Face released a benchmark called ThinkingBox. This benchmark evaluates agents based on the backend state and side effects they leave behind, rather than the sentences they generate, and tests whether they can repeat these actions twenty times. An example provided is an agent that successfully makes nine clean tool calls and closes a ticket as resolved, but fails when the carrier exception remains open, rendering the required end state "on hold."

A grader examining the tool calls sees nine well-formed calls, while the database indicates otherwise. The Witness Was the Suspect on Dev.to drew a parallel, arguing that when an agent writes its own success log, the incident record is authored by the incident's cause. The author outlines what they would require before trusting an agent's "done" status and what they would not claim.

They emphasize that status strings, tool-call traces, and a tidy final message, while useful for debugging, do not constitute a grade. The evaluation contract should assess the terminal state through a read path the agent cannot control and perform this assessment multiple times. This approach treats "done" when the state is red as a named failure rather than noise.

The author provides a menu of fields to answer, including the required end state, independent read path, the grader's role, the system of record, and done-while-red. They argue that consistent end states across multiple runs indicate a true capability, and not just a lucky run.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

More from Monday 5 October →