Do Audio Language Models Hear and Read Distinctive Features Alike?
Study finds most audio language models do not align heard and read features, except voicing in Qwen2.5-Omni.
Researchers test whether audio language models encode a phonological distinctive feature in the same direction when a phoneme is heard and when it is read. Across 6 models, 7 features, and 15 languages from 11 families, only voicing in the two Qwen2.5-Omni models exceeds a random-pairing reference after multiple-testing correction. In three of six models, voicing keeps one audio direction across 14 languages with enough minimal pairs. Model family, not size, predicts which stream represents a feature, and the random reference varies by a factor of seven between models.