Microsoft Research scales coding agents to 1,024 workers
A Microsoft Research paper says a leaderless multi-agent system raised a hard coding benchmark's pass rate from 34% to 55% as it scaled from one agent to 1,024.
Researchers from Microsoft Research described a system called Agensh in which up to 1,024 coding agents work in parallel with no central orchestrator. The agents coordinate instead through a shared workspace and a message channel, according to a paper posted to arXiv. Each agent gathers context, claims a sub-task, completes it and merges its result asynchronously, the paper said.
https://x.com/omarsar0/status/2104377054829613473
On the five hardest tasks in ProgramBench, the paper reported, scaling the team from one agent to 128 lifted the mean final test-pass rate from 19.31% to 28.78%. On a task involving the pandoc document converter, scaling to 1,024 agents raised the pass rate from 33.89% to 55.06%. Larger teams also reached a given pass rate sooner, according to the authors.
The paper was flagged by AI researcher Elvis Saravia on X and has not been peer-reviewed or reproduced by an outside lab. Its authors also reported that agents developed forms of cooperation that were not explicitly programmed and grew more common as team size increased, without detailing what those behaviors were.
Most production multi-agent systems today rely on some hierarchy or a central controller, and Agensh's shared-state approach is a departure the paper frames as a test of how far decentralized coordination can scale. The authors did not say whether the code or the benchmark harness would be released.