arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

像人类一样听?语音语言模型中的语音象征与感知对齐

Hearing Like Humans? Sound Symbolism and Perceptual Alignment in Speech Language Models

Yun-Shao Tsai, Chun-Wei Chen, Chee-En Yu, Yi-Cheng Lin, Hung-yi Lee

arXiv 2607.10162首次发表:更新:

发表机构

Graduate Institute of Communication Engineering National Taiwan University Taipei, Taiwan; Graduate Institute of Electrical Engineering National Taiwan University Taipei, Taiwan; Artificial Intelligence Center of Research Excellence National Taiwan University Taipei, Taiwan(; ; )

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究语音语言模型是否有语音象征倾向,通过真实人类语音录音对比模型与人类数据,发现模型听觉判断与人类感知对齐差,错过声学线索,开放权重模型表现不佳,弱点在语音表示方式。

AI 中文摘要

语音象征是人类将语音声音映射到诸如圆润或尖锐等感知质量的倾向,主要源于语音的声学而非拼写。语音语言模型(SLMs)是否有此倾向尚不清楚,因为之前的评估依赖文本或图像而非真实语音。我们使用真实人类语音录音进行研究,比较模型在听觉、跨模态和视觉方面的判断与人类数据。发现SLMs的听觉判断与人类感知对齐不佳,错过驱动人类直觉的声学线索,开放权重模型无法可靠地将听到的声音与相应形状联系起来。排除形状感知的视觉控制表明,弱点在于语音的表示方式,即感知对齐不取决于更强的视觉,而是取决于捕捉人类所听线索的语音表示。

英文摘要

Sound symbolism, the human tendency to map speech sounds to perceptual qualities such as roundness or sharpness, arises primarily from the acoustics of speech rather than spelling. Whether Speech Language Models (SLMs) share this tendency remains open, as prior evaluations rely on text or images rather than real speech. We study it using genuine human speech recordings, comparing model judgments against human data across the auditory, crossmodal, and visual components of the effect. We find that SLMs' auditory judgments align poorly with human perception and miss the acoustic cues, such as spectral tilt, that drive human intuitions, and open-weight models cannot reliably link a heard sound to its corresponding shape. With a visual-only control ruling out shape perception, the weakness localizes to how speech is represented, suggesting that perceptual alignment depends not on stronger vision but on speech representations that capture the cues humans hear.

CommentsSLT 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑