Do LLMs Act on What They Know? From Partner Representations to Cooperative Actions
Across eight LLMs, probes decode Hanabi partner conventions better than the models act on them.
The paper tests whether frozen LLMs convert decoded partner conventions into cooperative receiving decisions in a controlled Hanabi-derived setting with scripted hints. Across eight models, linear probes recovered intent conventions substantially more accurately than target conventions, yet choices often disagreed with the sender's convention. Stating rules produced modest, model-dependent gains, while externally computed action recommendations produced larger average improvements. In a Qwen3-8B case study, matched-state reversals showed much greater sensitivity to action recommendations than rule statements, and tested activation transfers did not reliably reproduce oracle benefits.
- Eight frozen LLMs were tested in a scripted Hanabi cooperation environment.
- Probes recovered intent conventions more accurately than target conventions.
- Action recommendations improved cooperation more than general rule statements.
- Qwen3-8B was more sensitive to action advice than to reversed rules.
Full article162 words · extracted from arxiv.org · click to collapse
Cooperation with unfamiliar partners requires adapting to communication conventions that are not known in advance. We study this problem in a controlled Hanabi-derived environment with scripted hint generation, LLM-controlled receiving decisions, and frozen model weights. Across eight LLMs, linear probes recover intent conventions substantially more accurately than target conventions, yet receiving choices do not consistently agree with the sender's convention. We compare probe-predicted and ground-truth conventions presented either as general rules or as externally computed action recommendations. Rule statements yield modest and model-dependent changes in cooperation, whereas action translation produces larger gains on average. In a Qwen3-8B case study, matched-state statement reversals reveal much greater sensitivity to action recommendations than to rule statements. Activation transfers from oracle-action and non-oracle hint-restatement donors improve intent accuracy on both action classes, but the tested alternatives do not reliably reproduce these benefits. Together, these results distinguish convention decodability, sensitivity to convention information, and cooperative performance, and highlight limitations in turning available partner information into receiving decisions.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.08129