arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

听见却未留意:音频-语言模型中的副语言信息编码与丢失

Heard but Not Heeded: Paralinguistic Information Encoding and Loss in Audio-Language Models

Bhuvan Koduru, Dareen Safar B Alharthi, Rita Singh, Bhiksha Raj

arXiv 2609.00727首次发表:更新:

发表机构

Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究分析四种开源音频-语言模型的副语言信息编码与丢失,发现模型虽在音频编码器顶层强编码说话风格,但信息在输出前退化,模型分内容、声学驱动两类,揭示了模型编码与使用信息的差距。

AI 中文摘要

音频-语言模型旨在理解语音,但目前尚不清楚它们是否能捕捉到除了所说内容之外的说话方式。我们对四种开源模型(Whisper-large-v2、Qwen2-Audio-7B Instruct、Qwen2.5-Omni-7B和Chroma-4B)中的副语言信息进行了机制分析,使用带有可控说话风格的Expresso数据集。我们结合中心核对齐、留一说话人评估的线性探测、开放式语调预测以及内容韵律泄漏指标,追踪风格信息如何从音频编码器传递到最终输出。所有模型在编码器后期(即音频编码器顶层三分之一的层)都对说话风格进行了强编码,但这些信息在到达输出前会持续退化。投影器重塑表示几何结构但不移除信息,而解码器根据架构和训练目标,在保留风格的程度上存在差异。在输出层面,模型分为两种行为:一些是内容驱动的,其预测主要依赖文本;另一些是声学驱动的,其预测随说话风格变化。泄漏指标量化了这种差异,定性结果也证实了这一点。总体而言,我们发现了模型编码内容与使用内容之间的差距,凸显了当前音频-语言模型的一个关键局限。

英文摘要

Audio language models are designed to understand speech, yet it remains unclear whether they capture how something is said beyond what is said. We present a mechanistic analysis of paralinguistic information in four open source models, Whisper-large-v2, Qwen2-Audio-7B Instruct, Qwen2.5-Omni-7B, and Chroma-4B, using the Expresso dataset with controlled speaking styles. We combine centered kernel alignment, linear probing with leave one speaker out evaluation, open ended tone prediction, and a content prosody leakage metric to trace how style information moves from the audio encoder to the final output. All models strongly encode speaking style in the late encoder, that is, the top third of the audio encoder's layers, but this information is consistently degraded before reaching the output. The projector reshapes representation geometry without removing information, while decoders differ in how much style they preserve depending on architecture and training objective. At the output level, models fall into two behaviors. Some are content driven, where predictions depend mainly on text. Others are acoustic driven, where predictions vary with speaking style. The leakage metric quantifies this difference, and qualitative results confirm it. Overall, we identify a gap between what models encode and what they use, highlighting a key limitation in current audio language models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑