Urgent.News

What's breaking now, across thousands of outlets.

AI

Presentation: Building Reusable Evaluation Frameworks for Agentic AI Products

Susan Chang explains how Elastic transitioned from siloed, ad-hoc AI agent evaluations to a unified, production-grade framework. She discusses balancing LLM-as-a-judge with deterministic rules, bridging Python data science evals with TypeScript production code, and implementing deep tracing to catch regressions across complex RAG and cybersecurity workloads while preserving domain context. By…

Susan Chang, a Principal Data Scientist at Elastic, discussed the transition from ad-hoc AI agent evaluations to a unified, production-grade framework during the QCon AI event. She explained their journey in building reusable evaluation frameworks for agentic AI products.

Elastic, known for tools like Elasticsearch and Kibana, has various AI agents built on top of their platform, serving diverse purposes such as cybersecurity, observability, and security analyst tools. To ensure quality and prevent regression, they implemented tracing and evaluation mechanisms.

Chang highlighted the importance of balancing LLM-as-a-judge with deterministic rules, as well as bridging Python data science evaluations with TypeScript production code. They implemented deep tracing to catch regressions across complex RAG and cybersecurity workloads while preserving domain context.

The company built reusable security analyst tools using AI agents on top of Elastic, enabling users to extract relevant logs and summarize them. This capability was demonstrated through an attack discovery agent that pulls logs from Elastic to identify potential cyber-attacks. The team also developed enterprise chatbots using proprietary data stored in Elastic.

To evaluate their AI agents, Chang's team created custom datasets and evaluation metrics tailored to each use case. For instance, in the attack discovery scenario, they used precision and recall to match alert IDs, while also enforcing factuality scores, avoiding MITRE tactic hallucinations, and measuring similarity scores.

Written by urgent.news from InfoQ's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at infoq.com →

More in AI

More from Monday 5 October →