{
  "id": 12139389,
  "title": "Presentation: Building Reusable Evaluation Frameworks for Agentic AI Products",
  "url": "https://urgent.news/2026/10/05/presentation-building-reusable-evaluation-frameworks-for-agentic-ai",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-05T11:19:00.000Z",
  "source": {
    "name": "InfoQ",
    "slug": "infoq",
    "url": "https://www.infoq.com/presentations/elastic-ai-agent-evaluations/"
  },
  "original_language": "en",
  "account": "Susan Chang, a Principal Data Scientist at Elastic, discussed the transition from ad-hoc AI agent evaluations to a unified, production-grade framework during the QCon AI event. She explained their journey in building reusable evaluation frameworks for agentic AI products.\n\nElastic, known for tools like Elasticsearch and Kibana, has various AI agents built on top of their platform, serving diverse purposes such as cybersecurity, observability, and security analyst tools. To ensure quality and prevent regression, they implemented tracing and evaluation mechanisms.\n\nChang highlighted the importance of balancing LLM-as-a-judge with deterministic rules, as well as bridging Python data science evaluations with TypeScript production code. They implemented deep tracing to catch regressions across complex RAG and cybersecurity workloads while preserving domain context.\n\nThe company built reusable security analyst tools using AI agents on top of Elastic, enabling users to extract relevant logs and summarize them. This capability was demonstrated through an attack discovery agent that pulls logs from Elastic to identify potential cyber-attacks. The team also developed enterprise chatbots using proprietary data stored in Elastic.\n\nTo evaluate their AI agents, Chang's team created custom datasets and evaluation metrics tailored to each use case. For instance, in the attack discovery scenario, they used precision and recall to match alert IDs, while also enforcing factuality scores, avoiding MITRE tactic hallucinations, and measuring similarity scores.",
  "summary": "Susan Chang explains how Elastic transitioned from siloed, ad-hoc AI agent evaluations to a unified, production-grade framework. She discusses balancing LLM-as-a-judge with deterministic rules, bridging Python data science evals with TypeScript production code, and implementing deep tracing to catch regressions across complex RAG and cybersecurity workloads while preserving domain context. By…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}