Verifiable Social Reasoning for LLM Assistants
Researchers introduce Fuse, a multi-agent simulation benchmark with hidden motives that evaluates LLM social reasoning across 12 models, finding that user mediation compounds difficulty and that models are systematically sensitive to biased user framing.
The paper introduces Fuse, a multi-agent simulation framework for evaluating LLM assistants' social reasoning in consultation settings. A target agent with a hidden motive interacts with other agents, including one representing the user, who then consults the evaluated assistant to infer the motive; this design provides verifiable ground truth by construction. Simulation faithfulness was validated through a human study with 24k annotations. Applied to 12 LLMs, the benchmark shows that user mediation compounds the inherent difficulty of social reasoning, that models are systematically sensitive to biased user framing, and that longer conversations do not reliably improve prediction accuracy; the arXiv report additionally notes that models may need more details than humans. The framework and a 21k-example dataset are open-sourced. The two source reports (Hugging Face daily papers, 2026-09-14; arXiv cs.AI/cs.LG/cs.CL, 2026-09-15) agree on all core findings.
- Fuse is a multi-agent simulation framework in which a target agent with a hidden motive interacts with an agent representing the user, who consults the evaluated assistant, yielding verifiable ground truth by construction.
- Simulation faithfulness was validated through a human study with 24k annotations.
- 12 LLMs were evaluated with the benchmark.
- User mediation compounds the inherent difficulty of social reasoning for LLMs.
- Models show systematic sensitivity to biased user framing.
- Per the arXiv report, models may need more details than humans.
- Longer conversations do not reliably (arXiv: 'not always') improve prediction accuracy.
- The Fuse framework and a 21k-example dataset are open-sourced.
Coverage timelineoldest first · each row is one article
- · 5d agoVerifiable Social Reasoning for LLM Assistants
Hugging Face daily papers· 30
Researchers introduce Fuse, a multi-agent simulation benchmark evaluating LLM social reasoning, showing models are overly sensitive to biased user framing.
- · 4d agoVerifiable Social Reasoning for LLM Assistants
arXiv cs.AI / cs.LG / cs.CL· 18
Fuse, a multi-agent simulation with hidden motives, evaluates LLM social reasoning, revealing compounding difficulty from user mediation and bias sensitivity.