Urgent.News

What's breaking now, across thousands of outlets.

AI

Import AI 466: The bitter lesson for robotics, AIs complete week-long programming tasks; and OpenAI's accidental AI hacker

The warning shots will continue until civilization wakes up

Import AI 466: The bitter lesson for robotics, AIs complete week-long programming tasks; and OpenAI's accidental AI hacker

Import AI 466 reports on recent AI advancements in programming and robotics. Epoch and METR released MirrorCode, a benchmark for testing AI's ability to complete long-horizon programming tasks. Leading AI models, such as Opus 4.7 and GPT-5.5, demonstrated impressive results, successfully reimplementing complex programs in various languages. However, eight out of 25 target programs were never solved to a 100% threshold, indicating that AI still faces challenges.

Meanwhile, Anthropic showcased how powerful general-purpose models can enhance robot capabilities. Anthropic's Opus 4.7 autonomously completed robot tasks 20 times faster than human records. While humans needed 181 minutes to complete all tasks, Opus 4.7 finished in just 9 minutes and 35 seconds. This progress suggests that as AI models improve, it could have flow-through benefits to robotics, making robots more capable and adaptable outside industrial environments.

Written by urgent.news from Import AI's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at importai.substack.com →

More in AI

More from Monday 27 July →