{
  "id": 6160661,
  "title": "AI agents are creating more work, not less — and OpenAI’s own numbers back it up",
  "url": "https://urgent.news/2026/09/07/ai-agents-are-creating-more-work-not-less-and-openais-own-numbers",
  "topic": "ai",
  "section": "AI",
  "published": "2026-09-07T19:05:17.000Z",
  "source": {
    "name": "The New Stack",
    "slug": "the-new-stack",
    "url": "https://thenewstack.io/openai-agent-research-bottleneck/"
  },
  "original_language": "en",
  "account": "OpenAI recently achieved its goal of deploying an \"automated research intern\" capable of handling well-defined tasks that would typically take researchers several days. Since then, the use of such AI agents has been steadily increasing throughout 2026, with agents logging 3.1 agent-workdays for every human workday by mid-August. However, the median researcher is spending more than $600 per day on inference at current API prices, with the 90th percentile spending over $7,000.\n\nThe concept of an agent-workday versus a human workday is crucial to understanding the impact of this technology. While agents can work faster, they do not necessarily produce proportional outputs. Researchers can run multiple agents simultaneously, which increases the overall workload but also necessitates more supervision. OpenAI defines an \"automated research intern\" as an agent that can complete well-defined research tasks requiring days of human effort, but a human still oversees the process.\n\nThe company has set a future goal of developing an \"automated AI researcher\" by March 2028. OpenAI has broken down the agents' work into six categories—Decide, Design, Build, Run, Analyze, and Communicate—and found activity increasing across all of them between January and August. However, agents only contribute minimally to deciding what research to pursue. Most of the work involves creating research and infrastructure code, monitoring experiments, and providing technical support, leading to a decrease in debugging office hours.\n\nDespite the rise in agent hours, this does not automatically translate to more useful research. OpenAI can track code output and experiment counts easily, but neither metric accurately reflects the progress made by the agents. Compute usage has also surged as the number of experiments has increased. OpenAI utilized a separate model to gauge how well agents performed on tasks of varying difficulty, finding that while success rates improved between January and July, humans had to intervene on more than half of successful tasks that would have taken a person four to eight hours.\n\nSecurity incidents have imposed limitations on Astra, OpenAI's persistent-agent model. Astra's capabilities allow researchers to hand off multi-day assignments, exacerbating the supervisory strain. On July 20, a series of outages caused by agents forced OpenAI to take its training container service offline, later restoring it with tighter restrictions. On August 7, the company tightened access again after early indications suggested Astra could reach the \"Critical\" cybersecurity threshold in its Preparedness Framework, restricting access to higher-security research areas and adding safeguards that developers may have already encountered as unexpected API interruptions.\n\nWorkloads have shifted between models rapidly in response to these restrictions. Astra-class GPU allocation fell by 59.2% the following week, but researchers moved much of the work to other models, which saw GPU allocation rise by 17.2%. This redistribution accounted for roughly 85% of the drop in Astra usage, demonstrating how easily workloads can be redirected when one part of the system is locked down.",
  "summary": "OpenAI says it hit a goal it set last fall, stating researchers are now using what the company calls an The post AI agents are creating more work, not less — and OpenAI’s own numbers back it up appeared first on The New Stack .",
  "key_points": [
    "AI agents log 3.1 agent-workdays per human workday by mid-August 2026",
    "Median researcher spends over $600 daily on inference at current API prices"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}