OpenAI quietly updates its evaluation metrics for GPT-6 Astra, making changes that appear to favor Astra and continuing to revise other metrics after launch (Emily Forlini/Fortune)
OpenAI has changed several evaluation benchmarks for its GPT-6 Astra model since first publishing a blog post announcement mid-afternoon on Sept. 3.
OpenAI has updated its evaluation metrics for the GPT-6 Astra model, making changes that appear to favor Astra. According to Fortune, several evaluation benchmarks for GPT-6 Astra have been changed since the model's launch was announced on September 3.
The GPT-6 Astra model has shown significant improvement across benchmark testing and includes new business features. However, concerns have been raised about its new 'recurrent depth' reasoning capabilities, which allow the model to consider a problem multiple times before taking an action. This is different from the standard chain-of-thought reasoning used in previous models.
OpenAI chief scientist Jakub Pachocki stated that the company will not accept degradation in its ability to monitor model alignment beyond a certain level and will withhold scaling until it can regain enough confidence. TechRadar reports that numerous experts have expressed concerns about the model's new reasoning architecture.
Brief written by urgent.news from Techmeme, TechRadar — 2 reports on this story. Machine-written — may contain errors; check the original before relying on it.
This story
This is one outlet's version. Read the fullest account.
- OpenAI quietly boosts some of Astra’s evaluation metrics, and continues to change others post-launch fortune.com
- Why is there so much worry about OpenAI Astra, and what issues could ‘recurrent depth’ reasoning cause? The experts weigh in techradar.com
- “Sorry for the messy rollout”: OpenAI launched GPT-6 Astra, but developers are locked out thenewstack.io
- OpenAI rolls out GPT-6 Astra to Pro customers on the $100/month or $200/month plans (Zac Hall/9to5Mac) 9to5mac.com
- Why OpenAI's GPT-6 Astra launch made CEO Sam Altman say sorry timesofindia.indiatimes.com
- OpenAI says it can't read all of Astra's reasoning and admits covert sandbagging would likely go uncaught, yet still calls it the world's most aligned model (Celia Ford/Transformer) transformernews.ai
- By declaring GPT-6 Astra to be AGI, OpenAI is being flippant and cementing the term's status as nothing more than marketing (M.G. Siegler/Spyglass) spyglass.org
- OpenAI unveils advanced GPT-6 Astra model dailytimes.com.pk