Urgent.News

What's breaking now, across thousands of outlets.

AI

OpenAI quietly updates its evaluation metrics for GPT-6 Astra, making changes that appear to favor Astra and continuing to revise other metrics after launch (Emily Forlini/Fortune)

OpenAI has changed several evaluation benchmarks for its GPT-6 Astra model since first publishing a blog post announcement mid-afternoon on Sept. 3.

OpenAI has updated its evaluation metrics for the GPT-6 Astra model, making changes that appear to favor Astra. According to Fortune, several evaluation benchmarks for GPT-6 Astra have been changed since the model's launch was announced on September 3.

The GPT-6 Astra model has shown significant improvement across benchmark testing and includes new business features. However, concerns have been raised about its new 'recurrent depth' reasoning capabilities, which allow the model to consider a problem multiple times before taking an action. This is different from the standard chain-of-thought reasoning used in previous models.

OpenAI chief scientist Jakub Pachocki stated that the company will not accept degradation in its ability to monitor model alignment beyond a certain level and will withhold scaling until it can regain enough confidence. TechRadar reports that numerous experts have expressed concerns about the model's new reasoning architecture.

Brief written by urgent.news from Techmeme, TechRadar — 2 reports on this story. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at fortune.com →

More in AI

Context Window Flooding: How Attackers Weaponize the Lost-in-the-Middle Attention Gap

An attacker does not need a clever jailbreak when they can make the model stop reading the system prompt through sheer volume.

  • Attackers manipulate context window to push system prompt into dead zone
  • Three context flooding attack types: padding, relevance flooding, tool result flooding
  • Agentic pipelines amplify attack surface through persistence and multi-agent orchestration

Working: Multi-Tenant Agent Isolation Failures: When One User's Context Bleeds Into Another's

Multi-Tenant Agent Isolation Failures: When One User's Context Bleeds Into Another's On March 20, 2023, a race condition in a Redis client library caused ChatGPT to return data across user boundaries. Payment information and chat history from one account appeared in a different user's session. The bug was not in the AI model.

  • Multi-tenant agent isolation failures occurred on March 20, 2023
  • Race condition in Redis client library leaked user context into others
  • Cryptographic namespace segregation per tenant required for mitigation

More from Sunday 6 September →