OpenAI Says It Has Reached Its 'Automated Research Intern' Milestone
OpenAI says its agents can now handle well-scoped research tasks lasting days, as agent runtime inside its research organization overtakes human labor.

Summary
OpenAI says it has reached the “automated research intern” milestone it set last fall: an AI system able to carry out well-defined research tasks under human direction, including work that would take a skilled researcher a few days.
The September 6 disclosure also gives a rare look at how much machine labor is already entering frontier-model research. By mid-August, OpenAI says its research organization was using 3.1 eight-hour agent-workdays for every human workday. The company is targeting an automated AI researcher by March 2028.
The claim is significant, but narrower than “AI scientist.” OpenAI says people still choose research priorities, judge which results matter, and decide whether to scale, pause, or deploy systems.
Quick Take
- OpenAI says its automated research intern milestone is now achieved.
- The system handles bounded research tasks that can take skilled humans days.
- OpenAI reports 3.1 agent-workdays per human workday, not 3.1× research output.
- Its next target is an automated AI researcher by March 2028.
What OpenAI says has changed
OpenAI describes a research workflow increasingly built around coding agents running in parallel. Researchers are delegating longer and more complex work, including research code, infrastructure tasks, troubleshooting, monitoring runs, and parts of experiment analysis.
The company says experiments per active experimenter reached an all-time high in August 2026. It also reports that agent use is moving beyond code generation toward broader parts of the research lifecycle.
Human steering remains substantial. OpenAI says more than half of successful tasks estimated at four to eight hours of human work still involved at least one human intervention.
Why it matters: AI is becoming part of the labor system that builds the next generation of AI, not just a product produced by that system.
Original-value milestone map
| Stage | Status | What remains human-led |
|---|---|---|
| Coding agents | Established | Task definition, review, research judgment |
| Concurrent research agents | Daily use | Coordination and prioritization |
| Automated research intern | OpenAI says achieved | Direction, scope, judgment, deployment decisions |
| Automated AI researcher | Target: March 2028 | Not yet achieved |
| Full recursive self-improvement | Unsolved | OpenAI says aligned, safe RSI remains an open problem |
This distinction prevents a misleading leap from “multi-day task completion” to “autonomous scientist.” The current milestone is about task horizon and useful execution under supervision.
What 3.1 agent-workdays actually means
The 3.1 figure is an operational measure of total agent runtime normalized to eight-hour workdays. It is not a direct measure of discoveries, useful experiments, or model improvement.
Three bottlenecks still matter:
- Human judgment: researchers decide what is worth pursuing.
- Intervention: longer tasks still frequently need human correction.
- Compute and integration: faster coding does not automatically accelerate every stage of AI research.
So the useful conclusion is not “OpenAI is 3.1 times more productive.” It is that machine labor has become larger than human labor by runtime inside the research organization, while humans remain the control layer.
The safety tension is now harder to ignore
OpenAI connects these trends to recursive self-improvement: AI systems helping accelerate the research that produces more capable AI systems.
At the same time, OpenAI says it does not yet know how to safely reach aligned, full recursive self-improvement. Its September 6 research post points to stronger controls and recent training pauses after the Hugging Face agent-security incident.
In a separate essay published the same day, OpenAI chief scientist Jakub Pachocki argues that no lab has solved alignment and monitoring well enough to keep scaling at maximum speed indefinitely. He says voluntary slowdowns should become common until shared safety thresholds exist.
That is not an announced development freeze. It is a warning from OpenAI's research leadership published at the same moment the company is documenting faster AI-assisted research.
AI World Scope take
The milestone matters less because of the word “intern” than because of the feedback loop it represents.
Frontier labs are starting to use AI as a meaningful production input for frontier AI research itself. If task horizons continue to grow while intervention rates fall, research automation could become a more important capability indicator than many conventional benchmark gains.
The next real threshold is not another increase in agent runtime. It is whether a system can choose promising research directions, execute experiments, interpret results, and iterate with much less human steering.
What to watch next
- Whether OpenAI publishes repeatable evaluations for the research-intern claim.
- Whether human intervention rates fall on multi-day research tasks.
- Whether research progress rises alongside agent runtime rather than merely agent usage.
- Whether the March 2028 automated-researcher target changes as safety constraints tighten.
Sources & Documentation
Sources & Documentation
Sources used for this article, with source type and publisher shown where available.
- officialResearch acceleration: The view inside OpenAIVisit Source
- officialAn Alien MindVisit Source
- reportingOpenAI’s chief scientist says no lab should keep scaling at maximum speedVisit Source