Urgent.News

What's breaking now, across thousands of outlets.

AI

The Agent Said It Was Done. The Database Disagreed.

Microsoft's ThinkingBox evaluates AI agents based on the records they generate and the changes they leave in databases, rather than solely on the sentences they produce. This benchmark assesses the consistency of AI agents across stateful business workflows when interacting with large language models (LLMs). The findings show that an agent can perform well in a single trial but still leave incorrect data in the database, leading to inconsistencies in overall performance.

Among the 12 LLM models tested, Claude Opus 5.5 performed the best at 67.16% in consistent performance, while Kimi-K3 had the broadest coverage, solving 93.89% of tasks at least once. However, Kimi-K3 had the lowest consistency, as only 13.41% of its tasks succeeded in all 20 attempts. The study concludes that pass@1 is an inadequate measure for deployment, as it does not account for the consistency of an agent's actions in real-world scenarios.

Written by urgent.news from Hugging Face's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at huggingface.co →

More in AI

More from Saturday 3 October →