Agentic reliability and evaluations : Enterprises that got burned by a bad eval are the most likely to remove humans from the loop, not the least
Across 108 enterprises, trust in automated agent evaluation rose sharply in July — and the failure rate it is supposed to predict did not move at all. The share of organizations that fully trust automated evaluation nearly tripled, from 5% in June to 13%, and the complaint that evaluations don’t match real-world outcomes fell 10 points. Yet the same share as last month — just under half — shipped…
Agentic reliability and evaluations: Enterprises that have suffered from a poor evaluation are the most likely to remove humans from the loop, not the least. Trust in automated agent evaluation surged in July, with the share of organizations fully trusting it nearly tripling from 5% to 13%. However, the failure rate that the evaluation is supposed to predict remained unchanged.
The failure is evident in the cross-tabs, showing that trust in automated evaluation is almost exclusively among those who have not yet been burned. Among those that have experienced a failure, only 4% trust automated evaluation, while among those that have not, 24% do. Interestingly, the trend of moving towards autonomy does not slow down; it accelerates after a burn.
This is the second wave of the VentureBeat Pulse Research agent reliability tracker, with July serving as the first reading on direction rather than position. What changed is the confidence in automated evaluation. In June, only 5% of enterprises fully trusted it, with the most common limitation being that evaluations poorly aligned with real-world outcomes (29%).
In July, 13% fully trust automated evaluation, and the alignment complaint has dropped to 19%, no longer the leading objection. However, the failure rate among organizations remains the same, with just under half deploying an agent or LLM feature that passed internal evaluations but caused customer-facing failures. The improvement in confidence is not due to better evaluations; trust is concentrated among enterprises that have not experienced a false-confidence failure.
The autonomy trajectory is flat, but the population inside it has shifted towards organizations with direct evidence that evaluations miss things. The vendor market is finally showing signs of settling, with specialist platforms gaining ground and switching intent cooling. Enterprises are now prioritizing fit over cost in selecting evaluation tools.
The sample for this research is senior and buyer-credible, with a slight shift towards retail/consumer industries in July compared to June. The sample size is large enough to support directional conclusions, but it should not be treated as a precise measurement.
Written by urgent.news from VentureBeat's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.