arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

音频语言模型在听觉与阅读时是否以相同方式表征区别性特征?

Do Audio Language Models Hear and Read Distinctive Features Alike?

Yuanhao Chen, Peter Chin

arXiv 2609.30167首次发表:更新:

发表机构

Thayer School of Engineering, Dartmouth College(达特茅斯学院塞耶工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究探究音频语言模型在语音和文本输入下是否以相同方向表征区别性特征,发现仅Qwen2.5-Omni模型的浊音特征显著,且模型家族而非规模决定模态差异。

AI 中文摘要

音频语言模型通过单一解码器处理语音和文本。我们探究该解码器在听到音素与阅读音素时,是否以相同方向表征同一区别性特征。针对仅在一个特征上存在差异的最小对立音素对,我们计算两个成员平均表征之间的偏移量。对这些偏移量取平均,得到每个模态(音频流和文本流)的方向,并测量两者之间的余弦相似度。由于两个模态在任意音素对上已达成一致,我们将每个度量与基于随机配对构建的参考值进行比较,而非与零比较。我们将此方法应用于来自11个语系的15种语言中的6个模型、7个特征。经过多重检验校正后,仅两个Qwen2.5-Omni模型中的浊音特征超过了该参考值,且参考值在不同模型间变化达七倍。在六个模型中的三个中,浊音在14种具有足够最小对立对的语言中表现出单一音频方向,其中两个模型的每一对语言均一致。预测哪个模态表征某一特征的是模型家族,而非模型规模。

英文摘要

Audio language models pass speech and text through a single decoder. We ask whether that decoder represents a distinctive feature in the same direction when a phoneme is heard and when it is read. For minimal pairs of phonemes differing in one feature, we take the offset between the two members' mean representations. Averaging those offsets gives a direction for each stream, and we measure the cosine between the two. Because the two streams already agree about arbitrary phoneme pairs, we compare every measure against a reference built from random pairings rather than against zero. We apply this to 6 models, 7 features and 15 languages from 11 families. Only voicing in the two Qwen2.5-Omni models exceeds that reference after correction for multiple testing, and the reference varies by a factor of seven between models. In three of the six models, voicing has one direction in audio across the 14 languages with enough minimal pairs to measure it, and every language pair agrees in two of them. The model family, not the model size, predicts which stream represents a feature.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑