i
News
News · 2026-10-11

Anthropic’s tests find more AI agents often buy speed, not quality

@neuronium_ai @neuronium_ai

Research: More AI agents spend tokens for barely better results

Cover: Anthropic’s tests find more AI agents often buy speed, not quality
Vals AI's benchmark shows cost per app versus score on Vibe Code Bench. Arrows point from each model's single agent to its team at the same reasoning effort. Teams cost far more but barely improve scores. | Image: Vals AI[

Vals AI's benchmark shows cost per app versus score on Vibe Code Bench. Arrows point from each model's single agent to its team at the same reasoning effort. Teams cost far more but barely improve scores. | Image: Vals AI[

Source: the-decoder.com

Anthropic’s tests suggest that adding agents often buys speed, not a meaningful improvement in quality. Larger teams reached a target score faster, but on one benchmark, moving from 10 agents to 100 produced only a small gain over 24 hours. The trade-off matters as companies build systems that spend more compute by dividing work among agents: the extra tokens may not earn a better answer.

The gains vary by task

In two internal tests using Opus 5.5, quality improved by less as teams grew. On some ProgramBench tests, faster results also required more tokens.

Fable 5.1 gained more noticeably when its team grew beyond 10 agents on a theorem-proving task in Lean, but it still scored below Opus 5.5 in every test. On a knowledge-base task, Fable’s result slipped slightly when the team grew from 30 agents to 100.

OpenAI researcher Noam Brown made a similar distinction on the Dwarkesh Podcast: multi-agent systems mostly improve speed rather than quality. Four agents completed tasks twice as fast, at twice the cost. At 16 agents, the ratio held, though efficiency fell slightly.

Brown said the payoff depends heavily on the task. Web research and mathematics can be parallelized; writing a novel cannot. Sending 10,000 agents to write one would be as pointless as sending 10,000 people. He added that very large agent teams remain little studied because testing them is expensive.

OpenAI developer Eric Provencher recently warned against large agent groups for a related reason: agents coordinate poorly. He called the resulting waste a “coordination tax.”

The cost of parallelism

The more interesting question is what the speedup is worth when quality barely moves. I think the results make agent count a poor proxy for capability: a system can finish sooner while consuming more tokens and delivering much the same answer.

That leaves a gap in the evidence. The tests show that huge teams can be costly and task-dependent, but do not establish when faster completion is valuable enough to justify the bill. Until that trade-off is clearer, scaling the team may be easier to measure than improving the work.

Source: the-decoder.com

Daily AI news

Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.

Only what matters — every day

Follow on X