Apple ML Research: multi-agent LLM teams hold back expert agents
Apple ML Research published a paper with an uncomfortable conclusion: multi-agent LLM teams without fixed roles do not achieve synergy and in fact hold back expert agents. The authors drew on organizational psychology and introduced the “strong synergy” metric — performance above that of the team’s best individual member. Self-organizing teams do not clear that bar, while structured approaches with fixed roles remain more effective.
AI-processed from Apple ML Research; edited by Hamidun News
Multi-agent LLM systems, in which agent models interact freely, without predefined roles or scripts, do not outperform a single well-calibrated expert — on the contrary, they hold it back. This is the main conclusion of a new paper by Apple ML Research, whose title speaks for itself: "Multi-Agent Teams Hold Experts Back."
What Apple ML Research studied
Multi-agent systems are one of the key tools of modern AI applications: several specialized agents work together — one gathers data, another analyzes it, a third formulates the answer. Most existing approaches ensure coordination through rigid structure — fixed roles, a set sequence of steps, or rules for aggregating results. This works, but it limits flexibility.
Apple ML Research asked what would happen if these constraints were removed and agents were allowed to interact completely freely. The authors call such a system a "self-organizing team" — coordination in it is not defined in advance but is supposed to emerge organically through the agents' mutual interaction. The question matters for the whole industry: modern agent-orchestration frameworks are currently moving precisely toward greater autonomy and free interaction.
Drawing on organizational psychology, the authors introduced a key metric — strong synergy: a team is considered synergistic only if it consistently outperforms the best of its individual members, not merely their averaged result.
Main finding: there is no strong synergy
According to Apple ML Research, self-organizing LLM teams do not achieve strong synergy. Moreover, under unrestricted interaction, the team holds back its expert members rather than amplifying them. Key findings of the study:
- Multi-agent teams without fixed roles do not outperform the best expert agent within their ranks.
- Free coordination limits, rather than unlocks, the potential of expert participants.
- Structured approaches with predefined roles and result-aggregation rules are consistently more effective than unstructured ones.
- The "wisdom of the crowd" effect, well documented in human groups, does not reproduce itself in LLM agents.
Why "wisdom of the crowd" doesn't work for LLMs
This is a break from how humans operate. In human groups without rigid hierarchy, mutual error correction and diversity of perspective often make it possible to find solutions unavailable to any single participant acting alone. Through the lens of organizational psychology, this is explained by the accumulated culture of collaboration and the shared context a group builds over time.
Language models do not accumulate that kind of context: every session starts over, without shared history or established norms of interaction among agents. As a result, free coordination produces noise rather than synergy — the expert agent within such a team performs worse than it would have acting alone.
What this means for developers of agentic systems
The practical takeaway of the study: if coordination is not explicitly built in — through roles, workflows, or result-aggregation rules — a multi-agent system risks losing to a single well-calibrated specialist. Adding agents is easy; organizing their interaction so it actually works is far harder.
Apple ML Research's work is a strong argument for the structured design of multi-agent systems: synergy does not arise from free interaction on its own — it must be engineered explicitly. Until language models learn to organically build the coordination that human groups build through culture and experience, self-organizing AI teams will keep losing out to structured ones — holding back, rather than amplifying, the best agents within their ranks.
What is "strong synergy" in the Apple ML Research study?
Strong synergy is a metric introduced by the study's authors based on organizational psychology: a team of LLM agents is considered synergistic only when it consistently outperforms the best of its individual members, not merely when it shows an above-average team result. This is precisely the threshold that self-organizing agent teams, according to Apple ML Research, fail to clear.
Why do self-organizing teams of LLM agents underperform solo experts?
Because language models lack what allows humans to coordinate effectively without rigid structure — an accumulated culture of collaboration and a shared context built up over time. Every session of LLM agents starts over without shared history or established norms of interaction, so free agent interaction produces noise rather than "wisdom of the crowd," and the expert agent within such a team performs worse than it would acting alone.
How can multi-agent system failure be avoided, according to
Apple ML Research?
Coordination needs to be built in explicitly — through fixed roles for agents, a defined sequence of steps (workflow), and specific rules for aggregating results. The study shows that precisely such structured approaches are consistently more effective than free, unrestricted agent interaction, which risks holding back rather than amplifying the best expert on the team.
Want to stop reading about AI and start using it?
AI News is a curated feed of AI/tech news. Hamidun Academy teaches you to use AI systematically in your work.
The AI world, distilled — once a week
Seven stories that actually mattered, hand-picked. No noise, no reposts, no press releases.
Done! Check your inbox for a confirmation.