The Missing Minimal Pair: Stereotype Evaluation in LLMs
Researchers propose dual minimal pairs and a mutual-information metric for more reliable LLM stereotype evaluation.
The paper argues that comparing log-likelihoods of a single pair of contrastive stereotype sentences is unreliable, because rewriting the same stereotype with an alternative attribute can produce logically inconsistent preferences. It proposes a dual minimal-pair evaluation with two comparison axes, plus a data-augmentation framework that generates paraphrases and alternate attributes for English, Russian, Spanish, and Chinese stereotypes. Two new metrics include a mutual-information measure between social groups and stereotyped attributes, intended for more stable aggregation across languages and models. Code is released at the authors' GitHub repository.
- Single contrastive sentence pairs can flip stereotype preferences after rewriting.
- A dual minimal-pair setup adds paraphrase and alternate-attribute comparisons.
- Data augmentation covers English, Russian, Spanish, and Chinese stereotypes.
- A mutual-information metric supports aggregation across languages and models.
Full article159 words · extracted from arxiv.org · click to collapse
A common approach to measuring bias in Large Language Models is to compare the log-likelihoods of two contrastive stereotype sentences. We argue that such single-pair comparisons are often unreliable: simply rewriting the same stereotype with an alternative attribute can yield logically inconsistent preferences. To address this, we propose a dual minimal pair setup that introduces two axes of comparison for robust stereotype evaluation. First, we present a data-augmentation framework that fills critical gaps in existing stereotype datasets by generating paraphrases and alternate attributes. We apply our framework on a set of English, Russian, Spanish and Chinese stereotypes. Second, we introduce two evaluation metrics tailored to the dual minimal pair setup. One of these metrics provides a new perspective on bias by modeling the mutual information (MI) between social groups and stereotyped attributes. This MI-based metric is better suited for aggregation and enables more robust comparisons of stereotype strength across different languages and models. Our code is available at https://github.com/stepanat/missing-minimal-pair/.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.08747