Urgent.News

What's breaking now, across thousands of outlets.

AI

Terminal-Bench-Science: Evaluating AI agents on scientific research workflows

Terminal-Bench-Science is a groundbreaking benchmark designed to evaluate the scientific capabilities of AI agents, with a focus on real-world research workflows. This innovative benchmark, led by researchers at Stanford University and built by a team behind Terminal-Bench, brings together domain experts from various scientific disciplines and institutions worldwide.

By measuring AI agent performance through a diverse set of challenging, expert-curated workflows, Terminal-Bench-Science aims to create a continuous feedback loop between scientific needs and AI development.

The first release of Terminal-Bench-Science includes 70 tasks spanning the life, physical, Earth, mathematical, and engineering sciences. These tasks cover a wide range of scientific endeavors, such as data analysis, statistical inference, simulation, optimization, theorem proving, image reconstruction, signal processing, inverse problems, sensor calibration, model fitting, classification, and scientific machine learning.

Researchers contributed to this benchmark through an open process on GitHub, with discussions and feedback taking place on the #tb-science channel on Discord.

To ensure the quality and relevance of the tasks, an extensive review process is in place. Approved proposals undergo scrutiny from domain reviewers, technical reviewers, and a bar raiser to confirm scientific validity, challenge, and feasibility. The 70 tasks selected for Terminal-Bench-Science 0.1 represent a challenging and selective set for AI agents to tackle, pushing the boundaries of frontier models.

Comparing AI agents' performance on Terminal-Bench-Science 0.1, Claude Opus 5 emerges as the top performer, achieving a 30% resolution rate. GPT-5.6 Sol with Codex, Claude Fable 5 with Claude Code, and GPT-5.6 Terra follow, with the latter two reaching higher resolution rates at a lower cost. Notably, GPT-5.6 Luna and Kimi K3 show limited capability, resolving less than 10% of tasks.

The cost-resolution plot illustrates the trade-offs between performance and computational resources, highlighting the diverse capabilities of different AI agents within the benchmark.

Written by urgent.news from Hacker News's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at terminal-bench-science.ai →

More in AI

Filling Silent Streams: How AI Avatars Keep Engagement Alive Without Viewer Comments

📝 Originally published (in Japanese) at forge.workstyle.tech . The Challenge of "Silence" in Unmanned AI Avatar Live Streams When creating a live stream where an AI avatar operates autonomously, the…

  • AI avatars struggle with silence in live streams
  • 75-second silence threshold prevents frozen avatars
  • Autonomous speech generated when no comments present

More from Friday 28 August →