ZeroHour
Story · 2 sources · 2 articlesfirst updated ()

Verifiable Social Reasoning for LLM Assistants

infoAI researchimportance 30
What's new: First merged summary for this story (no previous summary existed). Merged two reports covering the same paper; all core facts (Fuse design, 12 LLMs, 24k-annotation human study, bias sensitivity, 21k-example open-sourced dataset) agree across sources. The arXiv report adds one extra finding: models may need more details than humans. The two reports phrase the conversation-length result slightly…
Merged summary · glm-5.3-flash · rewritten as coverage arrives

Researchers introduce Fuse, a multi-agent simulation benchmark with hidden motives that evaluates LLM social reasoning across 12 models, finding that user mediation compounds difficulty and that models are systematically sensitive to biased user framing.

The paper introduces Fuse, a multi-agent simulation framework for evaluating LLM assistants' social reasoning in consultation settings. A target agent with a hidden motive interacts with other agents, including one representing the user, who then consults the evaluated assistant to infer the motive; this design provides verifiable ground truth by construction. Simulation faithfulness was validated through a human study with 24k annotations. Applied to 12 LLMs, the benchmark shows that user mediation compounds the inherent difficulty of social reasoning, that models are systematically sensitive to biased user framing, and that longer conversations do not reliably improve prediction accuracy; the arXiv report additionally notes that models may need more details than humans. The framework and a 21k-example dataset are open-sourced. The two source reports (Hugging Face daily papers, 2026-09-14; arXiv cs.AI/cs.LG/cs.CL, 2026-09-15) agree on all core findings.

  • Fuse is a multi-agent simulation framework in which a target agent with a hidden motive interacts with an agent representing the user, who consults the evaluated assistant, yielding verifiable ground truth by construction.
  • Simulation faithfulness was validated through a human study with 24k annotations.
  • 12 LLMs were evaluated with the benchmark.
  • User mediation compounds the inherent difficulty of social reasoning for LLMs.
  • Models show systematic sensitivity to biased user framing.
  • Per the arXiv report, models may need more details than humans.
  • Longer conversations do not reliably (arXiv: 'not always') improve prediction accuracy.
  • The Fuse framework and a 21k-example dataset are open-sourced.
ProductsFuse

Coverage timeline

  1. · 5d ago
    Hugging Face daily papers· 30
    Verifiable Social Reasoning for LLM Assistants

    Researchers introduce Fuse, a multi-agent simulation benchmark evaluating LLM social reasoning, showing models are overly sensitive to biased user framing.

  2. · 4d ago
    arXiv cs.AI / cs.LG / cs.CL· 18
    Verifiable Social Reasoning for LLM Assistants

    Fuse, a multi-agent simulation with hidden motives, evaluates LLM social reasoning, revealing compounding difficulty from user mediation and bias sensitivity.