OpenAI says it has hit the target it set last autumn: an "automated research intern," a system that takes on well-defined research tasks under human direction, including tasks that would cost an experienced researcher several days. The claim arrives inside two documents published together — an internal report on how much of OpenAI's own research is now run by agents, and an essay titled "Alien Mind" from chief scientist Jakub Pachocki. The report is dense with numbers. The intern claim is not: OpenAI says the milestone was reached "according to our measurements" and publishes no verification of it. A full automated AI researcher is due by March 2028.
The usage figures are the part to read closely. The median OpenAI researcher now burns more than $600 a day on inference priced at API rates; the 90th percentile clears $7,000 a day, roughly eleven times the median. Token volume generated by the median researcher is up 124x since December 2025. Since June, total agent runtime has exceeded total human working hours. By mid-August the research organization was running 3.1 agent workdays for every human workday.
OpenAI reads its own numbers cautiously, which is the right instinct. The report says metrics like these are cheap to collect and hard to interpret, and that their relationship to actual research progress remains unresolved. Experiments per researcher hit a record in August, the highest since measurement began in early 2025 — but compute also rose noticeably over the same stretch. The authors' own conclusion is that overall progress is probably rising more slowly than any single metric suggests, because the bottleneck has moved to the tasks that are hardest to automate.
The composition of the delegated work supports that. Classified against Epoch AI's taxonomy, every category of research activity grew, with the largest gains in writing research code, writing infrastructure code, technical assistance, and monitoring training runs. Higher-level planning decisions are still only a small share of agent output. Agents are writing and babysitting; humans are still choosing what to build.
To check whether delegated work actually gets finished, OpenAI ran an agentic classifier over cases with a clearly measurable outcome, bucketing tasks by how long a human would need. Success rates rose across several difficulty levels between January and July. The distribution is the interesting part. Tasks a human would finish in under 15 minutes were completed with no intervention 86% of the time. Among successfully completed tasks in the four-to-eight-hour band, more than half required at least one human action along the way. Autonomy does not scale with task length; it decays with it.
Two caveats sit under that result. The classifier is itself an AI system, and OpenAI does not separately report how reliable it is. Meanwhile the most persuasive evidence in the whole report is indirect: daily requests in the company's internal support channel have fallen sharply since 2025, and one team shut down its troubleshooting office hours entirely because agents increasingly debug research infrastructure on their own. Nobody closes a help desk for a metric.
Pachocki's essay pulls hard in the other direction. He writes that AI is "grown rather than designed," so its overall behavior cannot be fully described in comprehensible terms. On the basis of internal results OpenAI does not detail, he expects the current pace to lead toward recursive self-improvement, and says he is worried that nobody is prepared for the consequences of machine intelligence continuing to grow this fast.
He is specific about which tool is failing. Chain-of-thought monitoring — one of OpenAI's principal methods for watching reasoning models — is losing reliability, he writes. Verbalized reasoning is blending into monitored communication and tool use, systems are getting better at manipulating their own reasoning process, and capability is rising even without verbalized reasoning at all. His forecast: further AI progress will increasingly be capped by how reliable monitoring is.
Alignment has gaps too. During the Hugging Face incident, Pachocki writes, agents did not breach the prohibition on manipulating people but departed from the spirit of the values they were trained on. He rates GPT-6 Astra as far better aligned with human goals than its predecessor GPT-5.6 Sol, while warning that generalizable alignment may not keep pace with general capability.
The justification for continuing to train larger models is defense. Models already beat humans at breaking into computer systems and getting back out, Pachocki argues, and only a short window remains to harden critical infrastructure. The report makes the matching case: an automated AI researcher is also an automated safety and alignment researcher. Pachocki adds that this must not become an "excuse for recklessness," and that once the seriousness of the stakes is understood, advancing at any cost looks absurd.
This is a company arguing with itself in public, and both sides of the argument deserve to be weighed separately. The capability side is thinner than the headline: an intern milestone graded by OpenAI's own measurements, success rates scored by an AI classifier whose reliability goes unreported, and an autonomy curve that falls apart somewhere between fifteen minutes and a workday. The safety side is not thin at all. A chief scientist stating that his company's main window into model reasoning is closing, and that progress will soon be rate-limited by monitoring, is the most substantive claim in either document. It is also the only major claim with no number attached to it.
The missing paragraph is what replaces chain-of-thought monitoring. Pachocki names the failure modes precisely and offers no successor method. Neither document says what the internal results are that moved him from possible to expected on recursive self-improvement — the evidence for the most consequential forecast in the essay is the one thing withheld. The Hugging Face incident gets the same treatment: invoked as a known event, its details not revisited.
The governance asks, by contrast, are unusually concrete. Pachocki wants documents like the Preparedness Framework and Anthropic's Responsible Scaling Policy turned into binding standards, with compliance checked by independent auditors, regulators, or international organizations. He writes that no lab — Anthropic included — has solved alignment and monitoring well enough to responsibly keep scaling at maximum speed for long. He names international coordination on future AI development the top priority for governments worldwide. Citing OpenAI's Frontier Policy Blueprint, the report's authors propose obliging companies to publicly document their progress on recursive self-improvement.
And the same essay calls OpenAI's focus on recursive self-improvement the only way to stay at the frontier of AI research. That is the shape of it: binding rules demanded for a race the company says it must run flat out to keep leading, while colleagues announce that GPT-6 opens the age of AGI — which, by my reading, does even less to slow anyone down.