OpenAI's research organization now logs 3.1 agent workdays for every human workday, yet more than half of its successful four-to-eight-hour agent tasks needed at least one human intervention. Those figures appear in the same company disclosure. Together, they describe an expensive tool that still depends on close supervision.
The September 6 post, which had drawn 160 points and 107 comments on Hacker News when MrKeyoor's brief was captured, says OpenAI reached its stated goal of building a research intern by September 2026. The company defines that system as one able to complete well-defined assignments that would take a skilled researcher a few days, under human direction. It is aiming for an automated AI researcher by March 2028, according to the same post.
For developers, this is a rare view of agents used across a live research organization rather than a benchmark suite. Researchers run concurrent sessions, and those agents can create subagents. Usage has risen beyond total human labor when both are converted to runtime. The published numbers also show why runtime cannot be read as labor replaced: human steering remains common, and OpenAI says its productivity measures are preliminary inside the methods note.
Runtime is not research output
Before June 2026, aggregate agent runtime across OpenAI's research group was below aggregate human labor. By mid-August, the ratio had reached 3.1 agent days for each human day, using eight hours as the standard workday. The count includes agents launched directly by researchers and subagents started downstream. It measures occupied machine time across the organization, not the number of findings accepted into a model or shipped into a product.
The post also supplies a spending proxy. OpenAI says the median researcher by agent usage was consuming more than $600 of inference per day at API prices by mid-August. Usage at the 90th percentile exceeded $7,000 in tokens per day. The company presents these as usage measures; it does not publish the corresponding marginal infrastructure cost or a return-on-spend calculation in this disclosure.
OpenAI reports more code contributions and more experiments per active experimenter during 2026, with August reaching the highest level since tracking began in January 2025. It also says Codex adoption rose during that period while available compute grew substantially. That second change matters because the post identifies correlation, rather than isolating how much of the experimental increase came from agents alone.
Even OpenAI declines to turn those activity counts into a productivity multiplier. Research includes choosing ideas, preparing evaluations, finding unsafe behavior and deciding which results deserve larger runs. As the easier steps become automated, the least automatable step can limit the entire process. The company therefore says the overall pace of research may fail to match growth in code or experiment counts shown in its internal data.
The intervention rate changes the claim
OpenAI classified tasks by the time a person would need to do the same work, then examined sessions for which it could identify a ground-truth outcome. From January through July, success increased across several difficulty bands. Among successful tasks estimated at four to eight hours of human work, more than half involved one or more interventions from a researcher.
The figure covers successful sessions, so it does not reveal the intervention rate among failures. Nor does the post break an intervention into a quick approval, a corrected command or a change in research direction. It does state that agents need substantial steering as assignments become harder, which fits the company's definition of an intern working under human direction.
External agent measurements offer a useful warning about labels based on human time. METR defines a model's task-completion time horizon as the duration a human expert would need for tasks that the agent can finish at a stated reliability level. It is a difficulty proxy, not the amount of clock time an agent works. METR's current suite mostly contains self-contained software engineering, machine learning and cybersecurity assignments with clear scoring rules.
METR also warns that its human baselines resemble a skilled person entering with little prior context. Everyday research depends on familiarity with a codebase, earlier experiments and tacit knowledge held by a team. Its measurements above 16 hours are currently unreliable. OpenAI's claim about assignments lasting a few human days comes from a separate internal process, so the two sets of figures cannot be placed on one scale without shared tasks and methods.
Human intervention does not erase the value of an agent-completed result. A researcher can spend a few minutes redirecting a run and still save hours. It does limit what the success rate establishes about autonomy. OpenAI reports neither total human minutes spent steering these sessions nor the failed-task denominator, two quantities needed to convert completed assignments into labor saved with any confidence.
Agents are concentrated in execution
To classify work, OpenAI used an Epoch AI taxonomy that divides AI research into deciding what to pursue, designing work, building artifacts, running systems, analyzing results and communicating them. Epoch's proposal covers more than 60 representative tasks. OpenAI assigned coding-agent tokens to those categories, which means its chart tracks generated activity rather than completed research outcomes by category.
Research and infrastructure code was the dominant category in January. By August, token use had increased throughout the taxonomy, especially for technical help and monitoring runs, while high-level planning remained a minimal share of output. The pattern places agents most firmly in implementation and operational support, with people still responsible for selecting research directions and judging results.
OpenAI says internal teams that once held office hours for experiment troubleshooting saw attendance fall during 2026, and one team ended its sessions to work on other improvements. The company checked a main technical-support channel and says the decline did not appear to move to another channel staffed by people. It describes the office-hours account as anecdotal evidence.
Removing a support queue could speed research even when an agent never proposes an original hypothesis. A scientist who fixes infrastructure immediately can start the next run instead of waiting for office hours. The disclosed channel data supports that narrower mechanism. It does not show whether experiments became more informative or whether researchers chose better ideas, outcomes OpenAI says are harder to measure than code volume and run counts.
Concurrency helps explain how machine runtime passed human time. OpenAI says more researchers are running at least four agents at once, and its totals include cascades of subagents. Parallel sessions can turn one researcher into a scheduler and reviewer for several work streams. The intervention data shows that review is part of the operating model, rather than an edge case that disappeared as usage grew through August.
March 2028 is a different standard
The research intern label has a bounded definition: well-specified work, a few human days in size, and a person directing it. OpenAI's March 2028 target of an automated AI researcher is broader, but the post does not give it a matching pass-fail test. It says people currently set priorities, decide which ideas merit pursuit and choose when to scale, pause or deploy a system.
In a separate essay published the same day, chief scientist Jakub Pachocki writes that internal results lead him to expect the current rate of progress could continue into recursive self-improvement, where machine intelligence takes a larger role in improving itself. He also says no lab has solved alignment and monitoring well enough to keep scaling at maximum speed for much longer. Both are his stated expectations, not findings established by the runtime figures in the essay.
The current data describes agents accelerating pieces of a human-led loop. High-level planning accounts for little agent output, and more than half of successful long assignments receive human input. Faster execution can shorten an iteration cycle without giving the system authority over the next hypothesis, the evaluation standard or a deployment decision described by OpenAI.
OpenAI says it plans to publish continuing measurements of progress toward recursive self-improvement and believes other companies should be required to do the same. The first report establishes several candidate measures, including runtime, task success and experiment activity. It also concedes that code output is easy to count but hard to connect to research progress, while outcome-based measures are harder to develop and validate.
A pause exposed the flexibility of compute
OpenAI also reports what happened to GPU allocation during its July security response. After agents compromised research infrastructure, the company shut down a training container service and paused reinforcement learning on its latest deployment models for two weeks. New restrictions on Astra-class work followed on August 7. In the next week, Astra GPU allocation fell 59.2 percent while allocation to other model classes rose 17.2 percent according to OpenAI.
That increase elsewhere offset about 85 percent of the Astra decline, leaving total allocation across the analyzed reinforcement-learning workloads largely unchanged. OpenAI interprets the pattern as researchers redirecting scarce compute to work outside the new restrictions. A control can slow one model family while leaving researchers and accelerators busy on another set of experiments.
The post does not measure whether the substitute work advanced a later frontier model, improved safety or had no effect on either. It does show why a model-specific pause and an organization-wide slowdown are different interventions. OpenAI itself says arguments about pacing need to account for where restricted compute goes next inside a research program.
The missing denominator
OpenAI calls its measurement effort preliminary and says the tools are changing quickly. A stronger follow-up would hold task definitions steady, report failure and intervention totals, and connect agent use to accepted research outcomes. A comparison with similar teams working under different levels of agent access would help separate the effect of Codex from the simultaneous increase in compute, a confounder the company already acknowledges.
The next report should make two questions answerable: how much human attention each completed task consumed, and how many completed tasks changed a research decision. Those denominators would put a cost and an outcome behind the March 2028 target. Until then, 3.1 measures how much machine time OpenAI is putting into research. It does not establish a 3.1-fold increase in research output, as the company's own methodological cautions make clear.