Overwhelmed by Choice: Studying LLM Decision Making at Scale
LLM accuracy falls as candidate sets grow, and hierarchical or permutation inference recovers about 20 points at N=160.
Researchers systematically tested LLMs as the number of competing answer candidates increased and found substantial accuracy drops across tasks, prompts, and model scales. Long-context retrieval does not fully explain the loss. They describe gold-margin collapse, where confidence in the correct answer weakens, and a bias that makes early candidates hard to overturn. Hierarchical partitioning and permutation-based inference improved accuracy by about 20 percentage points at N=160 on HotpotQA and MIMIC.
- Accuracy degrades as candidate-set size grows across tasks and model scales.
- Gold-margin collapse shrinks the gap between the correct answer and strongest distractor.
- Later candidates exert weaker influence, so early preferences are hard to overturn.
- Hierarchical and permutation inference recover about 20 points at N=160.
Full article177 words · extracted from arxiv.org · click to collapse
Multiple-choice and candidate-selection evaluations are widely used to assess LLM reasoning and decision-making, yet most benchmarks contain relatively small candidate sets. It remains unclear whether conclusions drawn from these settings remain valid as the candidate space scales. We systematically evaluate LLMs as the number of competing candidates increases and find substantial accuracy degradation across tasks, prompting strategies, and model scales. Controlled analyses show that standard long-context retrieval explanations cannot fully account for this degradation. Instead, we identify two systematic failure patterns. First, gold-margin collapse: the score gap between the correct answer and the strongest distractor progressively shrinks, driven primarily by weakening confidence in the correct answer. Second, earlier candidate preferences become increasingly difficult to overturn, with later candidates exerting progressively weaker influence on the final prediction. Motivated by these findings, we evaluate hierarchical partitioning and permutation-based inference, which improve accuracy by roughly 20 percentage points at $N=160$ on both HotpotQA and MIMIC. Overall, our results identify candidate-set scale as an important evaluation-protocol variable and show that strong small-option performance does not necessarily imply robust large-scale candidate comparison.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.32809