Prompted Identity Degrades Cooperation in Multi-Agent LLM Systems
Revealing model-family labels makes multi-agent LLMs factionalize, raising cost and cutting cooperative success.
A study finds that exposing each agent's model family to its peers causes factionalism: agents prefer same-label partners even when the task does not reward that split. Tests used two cooperative games and a reasoning benchmark with nine to twenty-five agents from up to five open-weight families. Shuffled or arbitrary labels still create factions, while removing the label eliminates the behavior. Labeled groups spend about 30% more rounds and 55% more tokens, and success drops from 96% to 81%.
- Same-family labels cause unprompted clustering even without a reward for it.
- Shuffled or arbitrary labels still produce factions; removing labels stops it.
- Labeled groups need 30% more rounds and 55% more tokens.
- Cooperative success falls from 96% to 81%.
- Withholding identity labels is the reported mitigation.
Full article182 words · extracted from huggingface.co · click to collapse
Multi-agent LLM systems increasingly mix models from several providers, yet exposing each agent's underlying model identity to its peers significantly impairs cooperation. We show that when agents are aware of each other's model family, the group splits into clusters, where agents prefer interacting with others carrying their same label, although nothing in the task rewards or asks for such a split. We argue that the label itself causes this split, which we define as factionalism. We show and measure this phenomenon in two cooperative games and on a reasoning benchmark, with nine to twenty-five agents drawn from up to five open-weight model families. We further show that when the announced families are shuffled, or replaced by arbitrary labels, the factions still follow this information; when the label is removed, this behavior disappears. In strictly cooperative tasks, labeled groups spend on average 30% more rounds and 55% more tokens to reach a decision, and their success rate drops from 96% to 81%. The effect replicates across tasks, group sizes and model families. Withholding identity labels from the agents is simple and effective mitigation.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.35928