Research Shows AI Agent Teams Waste Tokens for Minimal Quality Gains
Reviews

Research Shows AI Agent Teams Waste Tokens for Minimal Quality Gains

TechNews Editorial
TechNews EditorialOct 11, 2026 · 2 min read
Share

Why it matters

This research matters because deploying multi-agent AI systems multiplies computational costs without guaranteeing meaningful performance improvements in most test scenarios.

The facts

  • Vals AI tested GPT-6 Sol and Claude Opus 5.5 as single agents and teams, finding teams cost up to 5.1x more with minimal quality gains.
  • OpenAI researcher Noam Brown noted that multi-agent systems buy speed rather than quality, with efficiency dropping at scale.
  • Experts warn that agent swarms risk wasted money due to high costs and coordination breakdowns.

Advertisement

AI agent teams deliver almost no better results than single agents according to recent research. Evals company Vals AI tested GPT-6 Sol and Claude Opus 5.5 on the Vibe Code Bench. Both models ran solo and as teams at medium and maximum reasoning effort levels.

Agent teams increase costs significantly

The team configurations cost between 1.8x and 5.1x more than single agents. Out of four comparisons between teams and solo agents, only one showed a statistically significant improvement. That improvement occurred with GPT-6 Sol at medium reasoning where the team scored 7.3 points higher. At maximum reasoning, the team setup gave neither Sol nor Opus 5.5 any real advantage.

The results suggest that the extra cost of agent teams isn't worth it in most cases. This is especially true when models are already running at full compute. Anthropic also saw quality gains shrink as it added more agents in two of its own tests with Opus 5.5.

Larger teams buy speed over quality

Larger teams reached a given performance level faster. However, going from ten to 100 agents only nudged scores up slightly after 24 hours. In separate ProgramBench tests, speed gains came with higher token usage.

A computing system runs increasingly large agent teams, with benchmark scores flattening while token consumption continues rising.
Illustration: AI & Tech News

On the knowledge base task, Opus 5.5 scored 0.53 with 1 agent, 0.70 with 10 agents, 0.71 with 30 agents, and 0.74 with 100 agents. On the Lean Theorem Proving task, scores were 0.39 with 1 agent, 0.66 with 10 agents, 0.66 with 30 agents, and 0.68 with 100 agents. Fable 5.1 showed stronger quality gains on the Lean theorem proving task above ten agents, but still scored below Opus 5.5 across all tests. On the knowledge base task, Fable's score actually dipped slightly when scaling from 30 to 100 agents.

Task type dictates agent efficiency

OpenAI researcher Noam Brown confirmed in the Dwarkesh Podcast that multi-agent systems mainly buy speed, not better quality. Four agents solved tasks twice as fast but also cost twice as much. At 16 agents, the pattern held but grew slightly less efficient.

The effect depends heavily on the task. Web research and math parallelize well, but writing a novel doesn't. Brown stated that throwing 10,000 agents at a novel would be just as pointless as throwing 10,000 people at it. He acknowledged that scaling to very large numbers of agents remains largely unexplored because the costs are simply too high. OpenAI developer Eric Provencher recently warned against using agent swarms for exactly this reason. He argued they are most likely wasted money because coordination between agents breaks down, calling it the coordination tax.

Newsletter

Get the best AI & tech news daily

A concise daily digest. Unsubscribe anytime.

We use your email only to send this newsletter.

Keep reading