More work did not mean more autonomy
The team used AI in 96.5% of the tasks it examined. Over four weeks, the median ratio of agent actions to human inputs rose from 11 to 28.5. The authors warn against reading that increase as evidence that agents became more autonomous: agents took more steps after each human decision, but did not necessarily make more decisions themselves.
Of 455 completed tasks assisted by AI, participants said 151 could not have been done without it, roughly one-third. Those tasks were spread across 27 of the 56 participants, rather than concentrated among a few especially active users. In these cases, AI did not simply speed up work already underway; it made work possible that otherwise would not have been attempted.
Atria Dawn Preview leads in AutomationBench, CyberGym, and MLE-Bench Lite but trails the field in GDPval and SWE-Bench Pro. | Image: Atria Team
Source: the-decoder.com
Over four weeks, the median number of agent actions per human input rose from 11 to 28.5. | Image: Atria Team
Source: the-decoder.com
People still chose the direction
The most common approach to choosing methods and parameters was “AI proposes, human chooses,” used in 55.4% of cases. Overall, people made 85.5% of decisions about methods and parameters, compared with 9.2% made by AI. People set the goals and scope of the work in 93.4% of cases.
Depending on the type of decision, AI proposals accounted for 17% to 55%. But its share of final decisions remained in the single digits. Even among the 151 tasks judged impossible without AI, a human chose the goal 95.4% of the time.
The same division showed up when tasks ran into problems. Of 588 tasks with a recorded difficulty, 76% continued after human intervention, while agents handled 23% on their own. Help usually meant providing context or refining requirements, in 35.2% of cases, or diagnosing the problem and changing the method, in 34.7%. People rarely took over the work directly: partial edits accounted for 3.2% of cases, and a full takeover for just 0.7%.
When AI output needed correction, the agent made changes itself in 75.4% of cases after receiving human feedback. The bottleneck was human judgment, not execution.
Participants said about a third of AI-assisted tasks couldn't have been completed without AI. | Image: Atria Team
Source: the-decoder.com
Agents often supply the method proposals, but humans make the final call in over 80 percent of cases. | Image: Atria Team
Source: the-decoder.com
The oversight problem
The authors describe three stages in AI’s role: first a subject of research, then a tool for individual tasks, and now a project partner that plans and revises its work within goals set by people. A possible fourth stage is recursive self-improvement, in which more powerful models build more powerful successors.
But doing well on the tasks a model was trained for does not mean it can develop its own successor. The open question, the researchers argue, is how AI could propose different research directions and judge their promise before results are available.
In three quarters of problem cases, human intervention moved work forward, mostly through context or diagnosis rather than taking over. | Image: Atria Team
Source: the-decoder.com
The Atria team describes AI's evolution from research object to project partner. The next stage, recursive self-improvement, remains an open question. | Image: Atria Team
Source: the-decoder.com
The study also points to a risk in how people supervise agents. When an agent’s work unfolds through a chain of steps too long for a person to review in full, oversight can become little more than formal approval. Many participants ran agents autonomously to avoid interrupting long tasks with repeated requests for confirmation. They chose that degree of autonomy for convenience, not after considering how much authority to delegate.
I think that distinction matters more than the rising action-to-input ratio. More work can move to agents while consequential choices stay with people—but only if those people can still understand what they are approving.
The study arrives amid debate over whether AI systems could begin improving themselves. Anthropic has argued that AI could build its own successor sooner than expected; CEO Dario Amodei has called for the industry to set a speed limit. Anthropic says people now make only a single-digit percentage of decisions about research direction at the company. OpenAI uses GPT-5.6 Sol across the development cycle, while Google and DeepMind let agents search for alternative strategies using recorded Dream-RSI search trajectories. Those agents improve the search strategy, not the model itself.
More than 1,000 employees at leading AI companies recently warned that their organizations may be close to automating AI research. A separate study by Princeton University and the UK AI Safety Institute reached a conclusion closer to Atria’s: advanced models can handle engineering tasks in research, but not important decisions that require judgment.
Daily AI news
Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.
Only what matters — every day
Follow on X