{
  "id": 5676647,
  "title": "Artificial Analysis Intelligence Index v4.2",
  "url": "https://urgent.news/2026/09/05/artificial-analysis-intelligence-index-v4-2",
  "topic": "ai",
  "section": "AI",
  "published": "2026-09-05T00:04:14.000Z",
  "source": {
    "name": "Hacker News",
    "slug": "hacker-news",
    "url": "https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-2"
  },
  "original_language": "en",
  "account": "The Artificial Analysis Intelligence Index v4.2 has been updated to present more intricate tasks and incorporate additional private test sets to deter gaming. This interim update, which follows the launch of Index v4 in January, has been delayed to maintain stability in light of recent advancements in the field.\n\nAmong the additions, AA-Briefcase, an internal evaluation featuring a private held-out test set, gauges models on realistic agentic knowledge work tasks within intricate, multi-week projects. It assesses models based on multi-task projects, thousands of source files, and grades them using rubric and pairwise methods for verifiable task success, analytical quality, and presentation quality.\n\nSurge AI has introduced GDP.pdf, a tool that evaluates single-turn professional document reasoning across 100 PDFs in ten domains, with models needing to synthesize information from 4,592 pages containing text, tables, charts, footnotes, and exclusions. Models are graded against 1,275 expert-authored criteria for the All-pass Rate, which requires every criterion to be met.\n\nWeighting adjustments now allocate 40% of the Index's total to private, held-out test sets, doubling the percentage from the previous version. This shift aims to lessen the potential for labs to manipulate evaluations. The grading infrastructure has also been enhanced, with improved prompts, error corrections, and stability in new model additions.",
  "summary": null,
  "key_points": [
    "AA-Briefcase evaluates agentic knowledge work tasks in multi-week projects",
    "GDP.pdf assesses professional document reasoning across 100 PDFs in ten domains",
    "Private test sets now account for 40% of the Index's total score"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}