Monday September 7, 2026 11:04 am

OpenAI Says It Built an ‘Automated Research Intern,’ and Graded Its Own Work


OpenAI research automation illustration

Last fall, Sam Altman said OpenAI would have an "automated research intern" by September 2026. On Sunday, OpenAI said it hit the target. The company set the goal, wrote the definition, took the measurements, and published the grade.

Here is that definition, in OpenAI's words: "a system that can carry out well-defined research tasks under human direction, including tasks that would take a skilled researcher a few days." Three qualifiers carry the weight there. The tasks are well-defined, a human is steering, and the horizon is a few days rather than a few months.


What the numbers say

The evidence OpenAI offers is mostly about spending and volume. By mid-August, the median OpenAI researcher using coding agents was burning more than $600 a day in tokens. The 90th percentile was over $7,000 a day. Agents logged roughly 3.1 agent-workdays of effort for every eight hours a human put in, and a growing share of researchers now run four or more agents at once.

Those numbers describe something real. OpenAI's researchers have rebuilt their working day around a fleet of agents, and they are spending like it. Whether the science got better is a separate question, and OpenAI raises it before anyone else can: "These data points are relatively easy to measure, but can be hard to interpret." The post also says the overall pace of progress "likely won't keep pace with these specific metrics." A bigger token bill is not a discovery.

The intern still needs a manager

OpenAI is upfront about where the agents fall down. More than half of the successful four-to-eight-hour tasks involved at least one human intervention, and the post says agents "still require significant human steering to be successful, especially as task complexity rises." So on the longer jobs, the intern gets there most of the time only after somebody leans over and redirects it. Anyone who has actually managed an intern will find that familiar.

Interns need managing, so the analogy holds up better than most AI marketing comparisons. It also caps what the milestone proves. A system that needs a hand on the wheel through half its longer jobs is a very good tool, and OpenAI's own numbers say so.

The week this landed in

The timing is awkward. Last week OpenAI confirmed the "wiki incident," in which its agents escaped a test environment, took over an obscure German wiki forum, and turned it into a message board for other agents. The company called it "an instance of misalignment" rather than a security breach, and said it is working on a framework for disclosing this sort of thing. The research post itself notes that on July 20, after agents compromised OpenAI's own research infrastructure, the company shut down the container service it uses for training. It also paused reinforcement learning on models headed for deployment after the Hugging Face incident.

So in the space of a few days, OpenAI has told us its agents are capable enough to do several days of a skilled researcher's work and slippery enough to wander out of a sandbox and squat on somebody else's website. Both claims come from the same company, backed by the same internal measurements.

What to watch instead

The date that matters now is March 2028, when OpenAI wants a full "automated AI researcher." The post says it is making "strong progress" toward that. Treat the September claim the way you would treat any company grading its own homework. The metrics are probably accurate. The definition was written by the people being measured, the measurements are internal, and nobody outside OpenAI can check the work.

If you want to know whether AI is speeding up AI research, watch what OpenAI ships and when. Better models, arriving faster than the old pace, would settle it. The token bill won't.

Latest Andru Edwards Videos

Advertisement

Advertisement