{
  "id": 11041939,
  "title": "Why Agent Observability Cannot Replace Evaluation",
  "url": "https://urgent.news/2026/09/30/why-agent-observability-cannot-replace-evaluation",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-09-30T18:16:38.000Z",
  "source": {
    "name": "HackerNoon",
    "slug": "hackernoon",
    "url": "https://hackernoon.com/your-agent-has-observability-it-doesnt-have-evals?source=rss"
  },
  "original_language": "en",
  "account": "The report, based on 1,340 responses collected between November and December 2025, reveals that while 57% of agents are in production (67% for teams with over 10,000 employees), only 89% of organizations have implemented observability across all agents, and just 62% include full per-step tracing. Additionally, only 52.4% engage in offline evaluations, and a mere 37.3% conduct online evaluations in production, with even lower numbers among production teams (44.8%). Despite the widespread adoption of observability, systematic evaluation of agent performance remains scarce, with quality factors such as accuracy, relevance, consistency, and tone cited as the top barriers preventing agents from being deployed in production. While teams can easily reconstruct what their agents did, they can only determine whether the outcome was satisfactory in approximately half of the cases. The OpenTelemetry GenAI semantic conventions define the standard attributes for monitoring LLM and agent workloads, but they do not include fields to measure correctness. Evaluating correctness requires reference to external data, business rules, or resulting states, which are not captured within the request/response cycle. Consequently, while observability metrics such as latency, token usage, and error rates remain constant regardless of the agent's performance, the true measure of an agent's effectiveness—its ability to produce correct outcomes—remains hidden.",
  "summary": "LangChain surveyed 1,340 practitioners: 89% have agent observability, 52.4% run offline evals, and fewer than a third do both. Traces tell you what an agent did",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}