Urgent.News

What's breaking now, across thousands of outlets.

AI

AI agents are creating more work, not less — and OpenAI’s own numbers back it up

OpenAI says it hit a goal it set last fall, stating researchers are now using what the company calls an The post AI agents are creating more work, not less — and OpenAI’s own numbers back it up appeared first on The New Stack .

AI agents are creating more work, not less — and OpenAI’s own numbers back it up

OpenAI recently achieved its goal of deploying an "automated research intern" capable of handling well-defined tasks that would typically take researchers several days. Since then, the use of such AI agents has been steadily increasing throughout 2026, with agents logging 3.1 agent-workdays for every human workday by mid-August. However, the median researcher is spending more than $600 per day on inference at current API prices, with the 90th percentile spending over $7,000.

The concept of an agent-workday versus a human workday is crucial to understanding the impact of this technology. While agents can work faster, they do not necessarily produce proportional outputs. Researchers can run multiple agents simultaneously, which increases the overall workload but also necessitates more supervision. OpenAI defines an "automated research intern" as an agent that can complete well-defined research tasks requiring days of human effort, but a human still oversees the process.

The company has set a future goal of developing an "automated AI researcher" by March 2028. OpenAI has broken down the agents' work into six categories—Decide, Design, Build, Run, Analyze, and Communicate—and found activity increasing across all of them between January and August. However, agents only contribute minimally to deciding what research to pursue.

Most of the work involves creating research and infrastructure code, monitoring experiments, and providing technical support, leading to a decrease in debugging office hours.

Despite the rise in agent hours, this does not automatically translate to more useful research. OpenAI can track code output and experiment counts easily, but neither metric accurately reflects the progress made by the agents. Compute usage has also surged as the number of experiments has increased. OpenAI utilized a separate model to gauge how well agents performed on tasks of varying difficulty, finding that while success rates improved between January and July, humans had to intervene on more than half of successful tasks that would have taken a person four to eight hours.

Security incidents have imposed limitations on Astra, OpenAI's persistent-agent model. Astra's capabilities allow researchers to hand off multi-day assignments, exacerbating the supervisory strain. On July 20, a series of outages caused by agents forced OpenAI to take its training container service offline, later restoring it with tighter restrictions.

On August 7, the company tightened access again after early indications suggested Astra could reach the "Critical" cybersecurity threshold in its Preparedness Framework, restricting access to higher-security research areas and adding safeguards that developers may have already encountered as unexpected API interruptions.

Workloads have shifted between models rapidly in response to these restrictions. Astra-class GPU allocation fell by 59.2% the following week, but researchers moved much of the work to other models, which saw GPU allocation rise by 17.2%. This redistribution accounted for roughly 85% of the drop in Astra usage, demonstrating how easily workloads can be redirected when one part of the system is locked down.

Written by urgent.news from The New Stack's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at thenewstack.io →

More in AI

More from Monday 7 September →