Is a Thousand Agents Better Than One?
When a task is too big for a single AI agent, the usual solution is to spin up several and assign each one a piece of the work. But that creates a new problem: who assigns the tasks, tracks progress, and brings the results together? Usually, that’s the lead agent. And the bigger the team, the harder it is for the lead to coordinate everyone.
Microsoft researchers have proposed a different approach. In Agensh, there is no lead agent: participants choose tasks for themselves, share progress, and combine their completed changes. In an experiment to recreate pandoc, a team of 1,024 agents passed more hidden tests than a single agent. The researchers also tested smaller teams. Across five challenging programming tasks, increasing the team from one agent to 128 raised the average share of tests passed by about 49%.
The idea is to scale not the language model itself, but the organization around it. The number of agents becomes a way to improve results in its own right, though it’s still unclear how well this approach works beyond programming.
What Takes the Place of a Lead Agent
In the usual setup, one agent acts as a manager: it breaks a large task into smaller ones, assigns them, and checks the results. But the manager can become a bottleneck. As the team grows, more effort goes into assigning work and pulling everyone’s results together.
Agensh removes this central point of control. Every participant gets the same overall goal and decides which subtask to take on. To keep the team from turning into a crowd, everyone has access to the same tools: a workspace, a messaging channel, and shared notes.
The workspace holds the code and work products. In the experiments, the researchers used version control: agents worked in separate branches, then proposed changes for inclusion in the shared version. The messaging channel helps them coordinate when tasks overlap or depend on one another. In the shared notes, agents post verified facts, approaches that didn’t work, current plans, and brief reports on their changes.
It’s like a development team without a manager handing out every assignment. Participants can see the shared plan and their colleagues’ results, pick up unclaimed work, and resolve conflicts directly.
In Agensh, participants choose tasks for themselves and share progress through a common workspace, messages, and notes.
How the Collaboration Loop Works
Every agent follows the same cycle, but works independently of the others. First, it reads the shared goal and checks what has already been done. Then it picks a subtask and records it in the shared notes. If it turns out someone else has already started the same work, the agents coordinate to split it up or adjust their plans.
The agent then works on the code, checks the result, and merges it into the shared version. If its changes conflict with a colleague’s work, the conflict has to be resolved first. Once finished, the agent reviews the team’s progress again and picks another task.
Everyone moves through this cycle asynchronously, so no one has to wait for the others to finish. Shared notes help prevent agents from repeating approaches that have already been tested. If an agent discovers that a solution doesn’t work, it can tell the team.
Each agent follows five steps: review progress, choose a subtask, complete it, check the result, and merge it.
This approach involves a trade-off. Autonomy eliminates the need to assign work to every agent in advance, but it doesn’t guarantee that everyone will choose a useful task. Agents may duplicate effort, run into conflicts when merging code, or misjudge the quality of one another’s work. Shared tools are intended to reduce these risks, but they don’t eliminate them automatically.
Testing the Approach on Large Programming Tasks
The researchers tested Agensh on five challenging tasks from the ProgramBench benchmark. Each task involved recreating a program’s behavior from scratch, with access to its compiled executable but no internet connection. The examples included FFmpeg, PHP, and pandoc. The agents had six hours to work.
Across the five tasks, the average share of tests passed rose from 19.31% with one agent to 20.68% with eight, 26.52% with 32, and 28.78% with 128. That last figure is an increase of 9.47 percentage points, or about 49% relative to the starting result. It’s a meaningful improvement, but the programs were still far from fully recreated.
The pandoc test showed a similar pattern at an even larger scale. One agent passed 33.89% of the tests, 128 passed 50.94%, and 1,024 passed 55.06%. Going from 128 to 1,024 agents added 4.12 percentage points. The team continued to improve its result, though the gain was much smaller than the jump from one agent to 128.
When recreating pandoc, the share of tests passed rose from 33.89% with one agent to 55.06% with 1,024.
More agents also helped teams reach intermediate results sooner. For example, on pandoc, the 128-agent team passed 30% of the tests within half an hour. Teams of 32 and eight reached that mark in an hour and an hour and a half, respectively. The single-agent run didn’t reach it in the first two hours.
If a team needs an acceptable result as quickly as possible, parallel work can help even without a guarantee of a complete solution.
On challenging tasks, larger teams generally reached the same test-passing rates sooner.
How the Team Learns to Organize Itself
The researchers looked not only at final scores, but also at records of how the agents worked together. With eight agents, they coordinated interfaces between parts of the program and avoided overlapping tasks. With 32, more effort went into checking and merging changes. For example, one agent might spot a bug in a colleague’s work, prompting the author to fix it before the team checked the result again.
In the 128-agent team, participants began to take on more consistent roles. Some checked changes; others helped merge code. Agents also reused collaboration strategies they had worked out. For example, an author might update and test their branch, then hand it to a colleague to review and merge.
In the 1,024-agent run, several participants specialized in merging changes. An agent could contact several potential helpers at once, choose the first one to respond, and cancel the other requests. If one participant couldn’t complete a task, another could pick up where it left off.
As the team grows, new ways of working together emerge—from coordinating subtasks to assigning roles.
The authors link these observations to the fact that all participants received the same prompt and had no predefined roles. In their interpretation, the ways of working together emerged as the team went along. But this is observed behavior from a single system working on program-recreation tasks—not proof that any large team of agents will organize itself effectively.
What We Still Need to Find Out
The results suggest that adding agents can improve the quality and speed of solving a difficult task. But the gains can’t be judged by test scores alone. A thousand participants require more computing power and generate more messages, changes, and potential conflicts. The paper focuses mainly on results achieved within the allotted time, not on the cost of achieving them.
The evaluation also has limits. The researchers tested a team of up to 1,024 agents only on pandoc, not on all five tasks. The experiment was also limited to programming and used a single base language model and one framework for running individual agents. It’s not yet clear whether the gains will hold in research, data analysis, or other tasks where results are harder to verify with automated tests.
The number of participants in a multi-agent system can improve results, provided they share a way to divide work, exchange knowledge, and combine changes that have been checked. The next question is how much that improvement costs—and whether it holds beyond programming. For a team of 1,024 agents, successful collaboration is already part of the task.
Read next
AgentKernel: How to Give an AI Agent More Freedom with Less Risk
How an AI agent selects past experience for a new task
JEV-as-a-Judge: The AI judge maintains 99% accuracy at a lower cost.
RRSI: How AI agent self-improvement enhances results and saves tokens
EvoOntology: Why is it difficult for AI agents to work with different
How AI Agents Can Save Context and Avoid Failures
How AI Simulates a User’s Thoughts
Coding agents skip looking at the app when the task gets long
AI covers six roles in game development, but skills rarely transfer
LLMs misread motives when the story comes through a biased user
Given a store for a year, the top-earning agent ranked 16th of 18 on fraud
Five levels of self-improving AI, and why level 5 barely exists
AI reviews in simple way
Every day we read fresh AI papers and retell the essentials in plain human language — no hype, no fluff. If you want to see where AI agents are heading before everyone else, subscribe.
New reviews — every day
Follow on X