Urgent.News

What's breaking now, across thousands of outlets.

AI

Artificial Analysis Intelligence Index v4.2

The Artificial Analysis Intelligence Index v4.2 has been updated to present more intricate tasks and incorporate additional private test sets to deter gaming. This interim update, which follows the launch of Index v4 in January, has been delayed to maintain stability in light of recent advancements in the field.

Among the additions, AA-Briefcase, an internal evaluation featuring a private held-out test set, gauges models on realistic agentic knowledge work tasks within intricate, multi-week projects. It assesses models based on multi-task projects, thousands of source files, and grades them using rubric and pairwise methods for verifiable task success, analytical quality, and presentation quality.

Surge AI has introduced GDP.pdf, a tool that evaluates single-turn professional document reasoning across 100 PDFs in ten domains, with models needing to synthesize information from 4,592 pages containing text, tables, charts, footnotes, and exclusions. Models are graded against 1,275 expert-authored criteria for the All-pass Rate, which requires every criterion to be met.

Weighting adjustments now allocate 40% of the Index's total to private, held-out test sets, doubling the percentage from the previous version. This shift aims to lessen the potential for labs to manipulate evaluations. The grading infrastructure has also been enhanced, with improved prompts, error corrections, and stability in new model additions.

Written by urgent.news from Hacker News's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at artificialanalysis.ai →

More in AI

Why AI Agents Keep Failing in Production (And What to Actually Build Instead)

Originally published on tamiz.pro . The Production Reality Check Every AI agent demo looks magical until it hits production.

  • AI agents fail in production due to hallucinations, context drift, and unpredictable failures.
  • Agents lack fallback mechanisms; incorrect LLM information is blindly processed, amplifying errors.
  • Shift focus from building agents to deterministic systems augmented by language models.

More from Saturday 5 September →