发表机构
University of Connecticut; The University of Texas at Austin(康涅狄格大学; 德克萨斯大学奥斯汀分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对音频语言模型(audio-LLM)韵律使用不足的问题,通过阶段特定探测阶梯与隐藏状态干预实验,发现模型能表征韵律却未在输出中表达,瓶颈在于韵律的使用而非感知。
AI 中文摘要
人类语音具有丰富的表达性,韵律除了承载词汇内容外,还传递语言和情感信息。因此,一个有能力的大型音频语言模型(audio-LLM)应支持对表达性语音的理解,不仅能转录所说内容,还能解读说话方式。然而,仅通过行为评估无法揭示模型在韵律输入上失败的原因,错误可能源于声学信息丢失、内部解释错误,或未能使用模型内部已有的表征。我们引入一种阶段特定探测阶梯,用于定位音频语言模型(audio-LLM)中的这些失败模式。在四个仅用于理解的音频语言模型(audio-LLM)中,韵律信息通常在音频路径中得以保留,并能在大型语言模型(LLM)的后期状态中解码,但仅部分在模型的最终响应中被表达。我们通过针对性的隐藏状态干预测试这种潜在表征的因果地位,每次干预都会按预期方向改变答案分布,且在大多数模型-任务组合中,仅在相关层进行一次编辑就足以推动模型朝向被抑制的韵律决策,不过这种恢复是定向的,而非选择性地恢复正确类别。特征级分析进一步表明,这种可恢复的信号可通过一个小子空间表达,分析中归因最高的部分特征与已知承载韵律信息的声学线索一致。在我们测试的匹配内容对比中,这些结果将反复出现的瓶颈定位在感知韵律之外,而在于使用韵律:能够感知并正确表征韵律线索的模型,仍可能无法在其答案中表达该线索。
英文摘要
Human speech is richly expressive, with prosody carrying linguistic and emotional information beyond the lexical content. A capable large audio-language model (audio-LLM) should therefore support expressive speech understanding, not only transcribing what was said but also interpreting how it was said. Yet behavioral evaluations alone cannot reveal why a model fails on prosodic input. An error may reflect loss of acoustic information, incorrect internal interpretation, or failure to use a representation that is already available inside the model. We introduce a stage-specific probe ladder for localizing these failure modes in audio-LLMs. Across four understanding-only audio-LLMs, prosodic information is usually preserved in the audio path and decodable in late LLM states. Yet it is only partially expressed in the model's final response. We test the causal status of this latent representation with targeted hidden-state interventions. Every intervention shifts the answer distribution in the predicted direction, and in most model--task cells a single edit at the relevant layer is sufficient to drive the model toward the suppressed prosodic decision, though this recovery is directional rather than a selective restoration of the correct class. Feature-level analysis further suggests that this recoverable signal can be expressed through a small subspace. Some of the highest-attribution features in this analysis align with acoustic cues known to carry prosodic information. Within the matched-content contrasts we test, these results locate the recurring bottleneck not in perceiving prosody but in using it. Models that hear and correctly represent a prosodic cue can still fail to express it in their answers.