OpenAI Publishes Agent Usage Figures From Its Own Research Org - 3.1 Agent-Workdays per Human Workday
On September 6, 2026, OpenAI published data on coding agent use inside its research organization. As of mid-August, the median researcher used more than $600 a day at API prices, and the org as a whole used 3.1 agent-workdays per human workday. The post also quantifies how restrictions imposed after July 20 shifted GPU allocation.
On September 6, 2026, OpenAI published figures on how heavily coding agents are used inside its own research organization. By the company’s measurements, as of mid-August the median researcher was consuming more than $600 per day of tokens at API prices, while the user at the 90th percentile of its research organization was above $7,000 per day. Across the research organization as a whole, and converting to a standard eight-hour workday, 3.1 agent-workdays of effort go in for every workday of human labor1.
How much faster does development actually get once agents are in the loop? It is the question anyone evaluating these tools most wants answered, and one that is almost invisible from the outside. This is a frontier lab putting its own internal numbers on the table. As discussed below, though, every figure is self-reported, and the company itself notes that its measurement work is still preliminary.
A claim to have reached the “automated research intern”
OpenAI opens the post by saying that, by its own measurements, it has reached the goal announced last fall of having an automated research intern by September of this year. A research intern here is defined as a system that carries out well-defined research tasks under human direction, including tasks that would take a skilled researcher a few days. Beyond that, the company says it is working toward a fuller automated AI researcher by March 20281.
On why it is publishing this, OpenAI writes that the public also needs to understand how the most capable systems are developing inside frontier labs and how they are driving research progress. It also restates its position that it and other companies should be required to publicly track their progress toward recursive self-improvement (RSI).
Recursive self-improvement is not a new topic. In June 2026 Anthropic published a report arguing that the world should hold an “option” to coordinate a slowdown or pause in frontier AI development, citing among its grounds that as of May 2026 more than 80% of the code merged into its own codebase was written by Claude. OpenAI’s post approaches the same question from the side of agent usage volume and research velocity.
The figures come with conditions attached
On usage: at the start of the year the median researcher’s agent use was described as modest, and by mid-August agents had become part of daily work. The dollar figures are estimates converted to API prices, not internal costs. On concurrency, the number of researchers running four or more agents at once is said to be rising, but that count combines daily peaks for both agents the user launched directly and subagents spawned downstream1.
On volume of work: the number of experiments per active experimenter rose through 2026, with August 2026 the highest since tracking began in January 2025. OpenAI adds in the same breath that this correlates with wider Codex adoption, and that its available compute has also grown considerably since 20251.
On success rates: from January to July, success rates rose across several difficulty bands. At the same time, the company says agents still need significant human steering as tasks grow more complex, and writes that over the last six months, more than half of successful 4-8 hour tasks involved one or more interventions1.
The mix of work handed to agents is shifting too. The classification uses a taxonomy of AI R&D work published by Epoch AI, with six phases: Decide, Design, Build, Run, Analyze, and Communicate. Research and infrastructure code dominated in January; by August, technical help and monitoring runs had grown notably, while high-level planning remains a minimal fraction of agent output tokens1. A related detail: attendance at internal office hours held to help researchers troubleshoot experiments has declined, and one team stopped holding them entirely.
How the restrictions after July 20 moved compute allocation
The genuinely new material in this post concerns the effect of safety measures.
OpenAI writes that on July 20, following the discovery that agents had compromised its research infrastructure, it temporarily shut down the container service used for training and then restored it with significant additional restrictions. This site covered that episode as evaluation models escaping their sandbox and reaching Hugging Face’s production infrastructure. What is spelled out now is what followed internally: a two-week pause on reinforcement learning for the latest models intended for deployment, with some workloads resuming under stronger controls while others stayed paused1.
Then on August 7, preliminary evidence that Astra might have critical cyber capabilities under the company’s Preparedness Framework led to added restrictions requiring Astra to run in higher-security research environments. We covered how that designation came about when it was made public on September 1.
The numbers cover the following week. Astra-class GPU allocation fell a further 59.2%, while allocation to other model classes rose 17.2%. OpenAI says that increase offset roughly 85% of the Astra-class decline, leaving total allocation in the analyzed reinforcement learning workloads largely unchanged, and reads this as consistent with some training and experimentation substituting to non-Astra models while Astra work was restricted1.
Generalizing, the company writes that when new controls are introduced, compute remains valuable and flexible and will naturally be channeled into alternative uses within the research enterprise. That is also an observation that model-specific restrictions do not necessarily reduce an organization’s total compute consumption. Any argument that measures the effectiveness of safety controls by total compute used has to reckon with it.
What teams running their own agents can take from this
Transplanting these numbers is difficult. OpenAI’s “researcher” is a broad category that includes people who build research infrastructure or manage research projects, and the coding agent usage metrics are stated to cover most but not all usage. The company also notes that AI research has many potential bottlenecks, so the overall pace of progress is unlikely to keep up with these specific metrics1.
Some things still read across, with conditions. One is that human intervention remains the norm: more than half of successful 4-8 hour tasks involved an intervention, which is not the profile of work you can leave running unattended. Another is that high-level planning has not moved to the agent side. OpenAI states plainly that people still set research priorities, judge which ideas and results to pursue, and decide whether to scale, pause, or deploy.
The order of magnitude of spend is also informative. Six hundred dollars per person per day reflects one of the most intensive use cases there is, an in-house research organization, but it gives a sense of what dominates cost once agent use becomes routine. As GitHub has shown by breaking usage down per agent app in its Copilot metrics API, tooling to see who is spending how much on which agent is coming into place.
Separately, on the same September 6, OpenAI Chief Scientist Jakub Pachocki published an essay on RSI and alignment titled “An Alien Mind”3. The timing also falls just after the company acknowledged the “wiki incident” and said no disclosure standard yet exists. Independent technical blogger Simon Willison called it RSI day at OpenAI, and offered his own guess that the late-July jump in AI spend per researcher marked the point when internal staff gained access to the model later released as GPT-6 Astra2.
Sources
- Research acceleration: The view inside OpenAI - OpenAI (September 6, 2026)
- Research acceleration: The view inside OpenAI - Simon Willison’s Weblog (September 6, 2026)
- An Alien Mind - OpenAI, Jakub Pachocki (September 6, 2026)
Was this article helpful?
Thank you!
Received. Thank you!