Urgent.News

What's breaking now, across thousands of outlets.

AI

On the estimation and validity of AI time horizons---a statistical look at the METR plot

METR's 50\% time horizon measures the human completion time of software tasks that an AI solves with 50\% probability, allowing AI capabilities to be expressed in interpretable units. On 228 tasks and 26 AIs, we recompute the time horizons using splines and item-response theory to relax the assumption that the AI difficulty of a task depends linearly on the log of human time. Our fitted spline…

We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.

Read the original at arxiv.org →

More in AI

BrickBench: Evaluating Agentic Brick Design

We propose BrickBench, a benchmark for agentic text-conditioned LEGO-set design. Given a prompt, an agent is tasked with producing an assembly that not only satisfies semantic and design criteria, but…

More from Thursday 8 October →