i
News
News · 2026-09-27

Atria’s AI agents did more model work, but people kept control

@neuronium_ai @neuronium_ai

Atria’s 744-billion-parameter Dawn Preview model was built with agents doing much of the work, from using tools to checking intermediate results against tests and other external signals. But the people involved still set the goals and made most decisions about methods and parameters. The project’s results point to a shift in how AI contributes to model development: agents can take on more steps, and even make some work possible, without taking over the judgment that directs it.

Cover: Atria’s AI agents did more model work, but people kept control

More work did not mean more autonomy

The team used AI in 96.5% of the tasks it examined. Over four weeks, the median ratio of agent actions to human inputs rose from 11 to 28.5. The authors warn against reading that increase as evidence that agents became more autonomous: agents took more steps after each human decision, but did not necessarily make more decisions themselves.

Of 455 completed tasks assisted by AI, participants said 151 could not have been done without it, roughly one-third. Those tasks were spread across 27 of the 56 participants, rather than concentrated among a few especially active users. In these cases, AI did not simply speed up work already underway; it made work possible that otherwise would not have been attempted.

Atria Dawn Preview leads in AutomationBench, CyberGym, and MLE-Bench Lite but trails the field in GDPval and SWE-Bench Pro. | Image: Atria Team

Atria Dawn Preview leads in AutomationBench, CyberGym, and MLE-Bench Lite but trails the field in GDPval and SWE-Bench Pro. | Image: Atria Team

Source: the-decoder.com

Over four weeks, the median number of agent actions per human input rose from 11 to 28.5. | Image: Atria Team

Over four weeks, the median number of agent actions per human input rose from 11 to 28.5. | Image: Atria Team

Source: the-decoder.com

People still chose the direction

The most common approach to choosing methods and parameters was “AI proposes, human chooses,” used in 55.4% of cases. Overall, people made 85.5% of decisions about methods and parameters, compared with 9.2% made by AI. People set the goals and scope of the work in 93.4% of cases.

Depending on the type of decision, AI proposals accounted for 17% to 55%. But its share of final decisions remained in the single digits. Even among the 151 tasks judged impossible without AI, a human chose the goal 95.4% of the time.

The same division showed up when tasks ran into problems. Of 588 tasks with a recorded difficulty, 76% continued after human intervention, while agents handled 23% on their own. Help usually meant providing context or refining requirements, in 35.2% of cases, or diagnosing the problem and changing the method, in 34.7%. People rarely took over the work directly: partial edits accounted for 3.2% of cases, and a full takeover for just 0.7%.

When AI output needed correction, the agent made changes itself in 75.4% of cases after receiving human feedback. The bottleneck was human judgment, not execution.

Participants said about a third of AI-assisted tasks couldn't have been completed without AI. | Image: Atria Team

Participants said about a third of AI-assisted tasks couldn't have been completed without AI. | Image: Atria Team

Source: the-decoder.com

Agents often supply the method proposals, but humans make the final call in over 80 percent of cases. | Image: Atria Team

Agents often supply the method proposals, but humans make the final call in over 80 percent of cases. | Image: Atria Team

Source: the-decoder.com

The oversight problem

The authors describe three stages in AI’s role: first a subject of research, then a tool for individual tasks, and now a project partner that plans and revises its work within goals set by people. A possible fourth stage is recursive self-improvement, in which more powerful models build more powerful successors.

But doing well on the tasks a model was trained for does not mean it can develop its own successor. The open question, the researchers argue, is how AI could propose different research directions and judge their promise before results are available.

In three quarters of problem cases, human intervention moved work forward, mostly through context or diagnosis rather than taking over. | Image: Atria Team

In three quarters of problem cases, human intervention moved work forward, mostly through context or diagnosis rather than taking over. | Image: Atria Team

Source: the-decoder.com

The Atria team describes AI's evolution from research object to project partner. The next stage, recursive self-improvement, remains an open question. | Image: Atria Team

The Atria team describes AI's evolution from research object to project partner. The next stage, recursive self-improvement, remains an open question. | Image: Atria Team

Source: the-decoder.com

The study also points to a risk in how people supervise agents. When an agent’s work unfolds through a chain of steps too long for a person to review in full, oversight can become little more than formal approval. Many participants ran agents autonomously to avoid interrupting long tasks with repeated requests for confirmation. They chose that degree of autonomy for convenience, not after considering how much authority to delegate.

I think that distinction matters more than the rising action-to-input ratio. More work can move to agents while consequential choices stay with people—but only if those people can still understand what they are approving.

The study arrives amid debate over whether AI systems could begin improving themselves. Anthropic has argued that AI could build its own successor sooner than expected; CEO Dario Amodei has called for the industry to set a speed limit. Anthropic says people now make only a single-digit percentage of decisions about research direction at the company. OpenAI uses GPT-5.6 Sol across the development cycle, while Google and DeepMind let agents search for alternative strategies using recorded Dream-RSI search trajectories. Those agents improve the search strategy, not the model itself.

More than 1,000 employees at leading AI companies recently warned that their organizations may be close to automating AI research. A separate study by Princeton University and the UK AI Safety Institute reached a conclusion closer to Atria’s: advanced models can handle engineering tasks in research, but not important decisions that require judgment.

Daily AI news

Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.

Only what matters — every day

Follow on X