ZeroHour
arXiv cs.AI / cs.LG / cs.CLpublished ()ingested Amir Taubenfeld

Verifiable Social Reasoning for LLM Assistants

infoAI researchimportance 18
AI summary · glm-5.3

Fuse, a multi-agent simulation with hidden motives, evaluates LLM social reasoning, revealing compounding difficulty from user mediation and bias sensitivity.

Fuse is a multi-agent simulation framework in which a target agent with a hidden motive interacts with other agents including one representing the user, who consults the evaluated assistant to infer the motive, providing verifiable ground truth by construction. Simulation faithfulness is validated through a human study with 24k annotations. Applied to 12 LLMs, it shows user mediation compounds social reasoning difficulty, models are systematically sensitive to biased user framing, models may need more details than humans, and longer conversations do not always improve performance. The framework and a 21k-example dataset are open-sourced.

  • Hidden-motive multi-agent simulation yields verifiable ground truth for social reasoning
  • Evaluated 12 LLMs; faithfulness validated with 24k human annotations
  • LLMs show systematic sensitivity to biased user framing
  • Longer conversations do not always improve prediction accuracy
  • Framework and 21k-example dataset open-sourced
Full article186 words · extracted from arxiv.org · click to collapse

LLM assistants are widely used for daily social advice, yet evaluating their social reasoning in such consultation settings remains challenging since (i) it requires setups where the assistant learns about social situations from subjective user narratives, and (ii) social properties, such as others' intentions, typically lack verifiable ground truth. To address these challenges, we introduce Fuse, a multi-agent simulation framework for studying user-mediated social reasoning. In Fuse, a target agent with a hidden motive interacts with other agents including one representing the user, who then consults the evaluated assistant to infer the target's motive, providing verifiable ground truth by construction. Simulation faithfulness is validated through a human study with 24k annotations. We apply Fuse to 12 LLMs and demonstrate its analytical utility by systematically isolating key factors, showing that (i) user mediation compounds the inherent difficulty of social reasoning; (ii) LLMs exhibit systematic sensitivity to biased user framing; (iii) models can require more details than humans need to reach a correct prediction; and (iv) longer conversations do not always improve performance despite providing opportunities for clarifying questions. We open-source Fuse and a dataset with 21k examples.

Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.17496