Urgent.News

What's breaking now, across thousands of outlets.

AI

GPT-6 Astra scores 95% on one robot task, 10% on another

Robocurve has run OpenAI's GPT-6 Astra and Anthropic's Claude Fable 5.1 against physical robot arms, and the results split sharply by task. Astra completed the easy task 19 times out of 20. On the harder task it managed 2 out of 20, exactly matching Fable 5.1. Robocurve published the Astra results on September 4, 2026, and the Fable 5.1 results the day before. Robocurve is an independent…

OpenAI's GPT-6 Astra and Anthropic's Claude Fable 5.1 were tested against physical robot arms by Robocurve, an independent evaluator. The results showed that GPT-6 Astra completed the easy task 95% of the time, and the harder task only 10%. In contrast, Claude Fable 5.1 achieved 40% success on the easy task and 10% on the harder task.

The evaluation was conducted on bimanual I2RT YAM arms, and each model ran 20 trials per task, with scores determined by human graders on a five-point scale. The block task cost $0.94 per run for GPT-6 Astra and $2.12 for Claude Fable 5.1, making Astra the cheaper option. The puzzle task cost $2.12 per run for both models, with 2 out of 20 correct answers for each.

The evaluation framework, called inspect-robots, is open-source and available on GitHub, providing transparency and reproducibility in the results.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Marketers face dilemma around rising bot and AI web traffic

Bot traffic is rising for e-commerce brands and interfering with retargeting strategies. But marketers are split on whether they should remain open to AI visitors or work harder to block.

  • Over half of web traffic now from bots and AI agents
  • Media companies struggle with scraped content for LLM tools
  • Some brands benefit from high-converting AI traffic

More from Tuesday 8 September →