{
  "id": 3855374,
  "title": "Terminal-Bench-Science: Evaluating AI agents on scientific research workflows",
  "url": "https://urgent.news/2026/08/28/terminal-bench-science-evaluating-ai-agents-on-scientific-research",
  "topic": "ai",
  "section": "AI",
  "published": "2026-08-28T00:06:51.000Z",
  "source": {
    "name": "Hacker News",
    "slug": "hacker-news",
    "url": "https://www.terminal-bench-science.ai/announcement"
  },
  "original_language": "en",
  "account": "Terminal-Bench-Science is a groundbreaking benchmark designed to evaluate the scientific capabilities of AI agents, with a focus on real-world research workflows. This innovative benchmark, led by researchers at Stanford University and built by a team behind Terminal-Bench, brings together domain experts from various scientific disciplines and institutions worldwide. By measuring AI agent performance through a diverse set of challenging, expert-curated workflows, Terminal-Bench-Science aims to create a continuous feedback loop between scientific needs and AI development.\n\nThe first release of Terminal-Bench-Science includes 70 tasks spanning the life, physical, Earth, mathematical, and engineering sciences. These tasks cover a wide range of scientific endeavors, such as data analysis, statistical inference, simulation, optimization, theorem proving, image reconstruction, signal processing, inverse problems, sensor calibration, model fitting, classification, and scientific machine learning. Researchers contributed to this benchmark through an open process on GitHub, with discussions and feedback taking place on the #tb-science channel on Discord.\n\nTo ensure the quality and relevance of the tasks, an extensive review process is in place. Approved proposals undergo scrutiny from domain reviewers, technical reviewers, and a bar raiser to confirm scientific validity, challenge, and feasibility. The 70 tasks selected for Terminal-Bench-Science 0.1 represent a challenging and selective set for AI agents to tackle, pushing the boundaries of frontier models.\n\nComparing AI agents' performance on Terminal-Bench-Science 0.1, Claude Opus 5 emerges as the top performer, achieving a 30% resolution rate. GPT-5.6 Sol with Codex, Claude Fable 5 with Claude Code, and GPT-5.6 Terra follow, with the latter two reaching higher resolution rates at a lower cost. Notably, GPT-5.6 Luna and Kimi K3 show limited capability, resolving less than 10% of tasks. The cost-resolution plot illustrates the trade-offs between performance and computational resources, highlighting the diverse capabilities of different AI agents within the benchmark.",
  "summary": null,
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}