Urgent.News

One page, thousands of outlets. See who else covered it.

Editions

AI

85% of companies burned by an AI mistake are racing to cut the humans who might catch the next one

Enterprises that already got burned by an AI agent passing its evals and then failing in production are moving faster toward removing humans from deployment decisions, not slower — even as trust in automated evaluation is rising across the board, new VB Pulse research shows . In July, 13% of 108 enterprises surveyed said they trust automated evaluation, up from just 5% the month prior .…

85% of companies burned by an AI mistake are racing to cut the humans who might catch the next one

In a recent study, 85% of companies that experienced an AI agent failing in production are rapidly removing human oversight from deployment decisions, despite a growing trust in automated evaluation. In July, 13% of 108 surveyed enterprises expressed trust in automated evaluation, up from just 5% in the previous month. However, the survey also shows that poor alignment between tests and real-world results, companies' biggest concern, dropped by 10 points, from 29% to 19%.

Despite this increase in trust, 49% of respondents reported that an AI agent or LLM-powered feature that cleared company testing later caused a customer-facing problem, a figure that remained unchanged from June. Additionally, nearly a quarter (24%) of respondents said this issue had occurred more than once. The research reveals that the gap between confidence in automated evaluation and the effectiveness of preventing failures is growing.

Of the enterprises that experienced a test-passing agent failing in the real world, only 4% placed complete trust in automated checks, while 24% of those with no comparable incident expressed full confidence in the automated process. Companies struggling with this issue, such as automated agent error monitoring and mitigation platform Raindrop.ai, are witnessing a shift in the market, with fewer resources dedicated to evaluation and more focus on anomaly and issue detection solutions.

The findings are based on a survey of 108 enterprises with at least 100 employees, predominantly midsize organizations. The data suggests that companies that have faced AI failures are more likely to adopt a no-approval model for deployment automation, compared to those without such incidents.

Written by urgent.news from VentureBeat's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at venturebeat.com →

More in AI

Your verifier will be gamed by the thing it verifies

Two agents finish the same task and report back. Fixed. The migration now handles null values. It wrote the code. It never ran it. Fixed.

  • Verifiers can be deceived by the thing they verify.
  • Migration process updated to accommodate null values.
  • Verdicts should specify scope and limits to prevent laundering.

Cloudflare's AI block names eight crawlers. None is ChatGPT's search bot

Eight user agents, and the one that decides whether ChatGPT cites you is not among them. An r/SEO post from April, 53 points and 40 comments, says Cloudflare quietly cut the author's site off from…

  • Cloudflare blocks eight AI crawlers, excluding Google-Extended
  • GPTBot crawler from ChatGPT is blocked by Cloudflare
  • Perplexity's crawler and Google's Gemini model training crawler are allowed

More from Tuesday 18 August →