{
  "id": 6361575,
  "title": "EVA-Bench Data 2.0 Benchmark Release: Covering 3 Domains, 121 Tools, and 213 Test Scenarios",
  "url": "https://urgent.news/2026/09/09/eva-bench-data-2-0-benchmark-release-covering-3-domains-121-tools-and",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-09-09T01:00:28.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/judy_miranttie/eva-bench-data-20-benchmark-release-covering-3-domains-121-tools-and-213-test-scenarios-26oh"
  },
  "original_language": "en",
  "account": "ServiceNow's AI research team has unveiled EVA-Bench Data 2.0, an enterprise-grade benchmark tailored for voice agents. This update significantly broadens the scope from a single domain to three major enterprise scenarios: airline customer service management, enterprise IT service management, and healthcare human resources service delivery. The expanded benchmark now encompasses 213 evaluation scenarios and 121 tools, nearly quadrupling the coverage of its predecessor. Each domain presents a unique set of scenarios: airline with 50, ITSM with 80, and HRSD with 83. The focus of EVA-Bench Data 2.0 is on real-world voice scenarios, with datasets meticulously crafted from genuine phone-based customer service workflows and modeled after production API specifications. The HRSD domain goes even further, integrating real US healthcare policies such as NPI provider identifiers, FMLA family leave regulations, and insurance coverage rules, ensuring the scenarios reflect day-to-day professional challenges. All 213 scenarios underwent cross-validation by three leading models - OpenAI's GPT-5.4, Google's Gemini 3.1 Pro, and Anthropic's Claude Opus 4.6, maintaining the benchmark's rigor while ensuring fairness and reliability. All datasets are freely accessible via Hugging Face Datasets. Plans are underway for a multilingual expansion to broaden the benchmark beyond its current English-only constraints. The comprehensive documentation of design principles and the generation process offers a robust implementation reference for those aiming to create their own evaluation datasets. This move by ServiceNow signifies a shift from theoretical discussions to practical deployment, emphasizing the importance of basing enterprise AI voice evaluation on actual business workflows. The multi-model validation process ensures the benchmark remains challenging yet fair across different model families. The healthcare scenarios, complete with intricate regulatory details, underscore the necessity of a high density of business-level information in evaluations for regulated sectors. The open-source release of scenario generation files serves as an excellent starting point for others designing enterprise AI evaluation datasets, providing a practical framework to identify production gaps effectively.",
  "summary": "This article is a deep-dive from JudyAI Lab — an AI engineering playbook series with 100+ published guides, 5,000+ weekly readers across 60+ countries, focused on the practical side of running AI agents, trading systems, and content pipelines in production. 📰 Key Summary ServiceNow AI's research team has released EVA-Bench Data 2.0, an enterprise-grade benchmark designed specifically for voice…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}