Urgent.News

What's breaking now, across thousands of outlets.

AI

OpenAI quietly boosts some of Astra’s evaluation metrics, and continues to change others post-launch

The numbers on OpenAI's blog about Astra kept changing. In some cases, the changes made Astra look better and rival models look worse.

OpenAI quietly boosts some of Astra’s evaluation metrics, and continues to change others post-launch

OpenAI has modified several evaluation benchmarks for its GPT-6 Astra model since announcing it on September 3. In some instances, the updated versions showed Astra performing better, while scores for rival models from OpenAI's competitor Anthropic worsened. The changes happened during an unusual rollout of the blog post, which was initially planned to go live at 2 p.m. ET but took almost two hours to become widely viewable online.

When OpenAI's X account tweeted the blog post at 3:32 p.m., the link initially returned an error message. It wasn't until 3:50 p.m. that the link was finally accessible, though some users still encountered the same error. OpenAI had published the blog post shortly after 2 p.m. but later retracted it for undisclosed reasons, which the company attributed to unrelated factors like a bug in the content management system or an internet outage.

Upon republishing, the blog featured different evaluation metrics that seemed to favor Astra. Some figures have continued to change even since then.

One notable change was Astra's reported hallucination rate, which dropped from 4.2% in the first internet archive snapshot of the blog post to 2% in the sixth snapshot taken at 5:20 p.m. The scores for Astra's predecessor, GPT-5.6 Sol, also decreased from 12.2% to 9.4%. However, OpenAI has since reversed these changes, with the hallucination rates returning to their original values of 4.2% and 12.2%.

OpenAI has also boosted GPT-5.6 Sol's internal version of the ExploitBench cybersecurity evaluation from 5.5% to 11.5%. The company stated that the 11.5% result reflects a reasoning level that is not commercially available for Sol. Astra excels at mathematics, a quality OpenAI highlights in the announcement page. While this metric did not change in the snapshots, OpenAI briefly altered the scores for GPT-5.6 Sol and Anthropic's latest model, Fable 5.1, making Astra appear significantly better at math than those models.

Written by urgent.news from Fortune's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at fortune.com →

More in AI

Why AI Agents Keep Failing in Production (And What to Actually Build Instead)

Originally published on tamiz.pro . The Production Reality Check Every AI agent demo looks magical until it hits production.

  • AI agents fail in production due to hallucinations, context drift, and unpredictable failures.
  • Agents lack fallback mechanisms; incorrect LLM information is blindly processed, amplifying errors.
  • Shift focus from building agents to deterministic systems augmented by language models.

More from Saturday 5 September →