OpenAI engineer: multi-agent swarms burn tokens without improving results

Eric Provencher, a developer on OpenAI's Codex, argues that running more than two parallel sub-agents typically consumes excessive tokens without boosting output quality, because agents distrust each other and repeatedly verify work. He points to a case where 1,393 agents spent $20,000 on a single Python refactoring that one agent could have completed at far lower cost. Provencher recommends delegating tasks to separate threads that notify the main agent only upon completion, rather than constant status polling.
Provencher's warning centers on what he calls a "coordination tax" in multi-agent workflows. When sub-agents don't trust each other's output, they repeatedly verify completed work, multiplying token consumption. He cites a striking example: 1,393 Fable agents spent $20,000 refactoring a single Python file, a task one Astra agent could have handled at a fraction of that cost.
The inefficiency stems partly from redundant system prompt execution across sub-agents and missing context that forces rework. Provencher suggests delegating tasks to separate threads that notify the main agent only upon completion, rather than constant status polling. He acknowledges OpenAI has yet to ship better solutions.
This warning could reshape how organizations approach agentic AI deployments. Companies investing heavily in multi-agent systems may reconsider their architecture choices, potentially favoring simpler, single-agent approaches for cost-sensitive tasks. The $20,000 example highlights real financial risks for businesses scaling AI workflows. However, Provencher's perspective comes from one developer's experience, and multi-agent systems may still prove valuable for complex, parallelizable problems. The broader impact could be a more measured, cost-conscious approach to AI agent adoption, with organizations weighing token efficiency against task complexity before scaling agent counts.