声音去向何方?追踪音频条件大语言模型中的声学信息损失
Where Does the Sound Go? Tracing Acoustic Information Loss in Audio-Conditioned LLMs
查看机构详情
- Seoul National University(首尔大学)
- KAIST(韩国科学技术院)
- NAVER Cloud(NAVER云)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文通过对比不同前端编码器并追踪信息流,发现音频条件大语言模型对声学信息利用不足主要源于读出对齐瓶颈,而非编码器信息丢失。
中文摘要 AI 辅助
音频条件语言模型往往未能充分利用韵律、情感和非语音声音等声学线索,这引发了一个问题:由ASR监督的前端是否在信息到达语言模型之前就将其丢弃。我们通过比较Whisper-Tiny和Whisper-Small与EnCodec、DAC-VAE和WavTokenizer在共享的Qwen3.5-4B音频语言模型流水线中的表现,来测试前端是否应为此负责,任务涵盖ASR、情感识别和声音描述。仅替换编码器并不能解决这种利用不足的问题:Whisper变体在整体上仍然最强,包括在情感和环境声音描述任务上。为了定位失败原因,我们追踪任务相关信息通过编码器、投影器、语言模型层和语言模型头的过程。线性探针和几何分析表明,判别性声学结构在最终语言模型层仍然可恢复,即使MCQA准确率落后于探针准确率多达83个百分点。由于答案格式和解码过程受到控制,这种任务依赖的差距指向内容特定的读出失败,而非通用格式偏差。LogitLens分析和针对性的语言模型头干预支持以下结论:声学利用不足不能仅由编码器侧信息损失解释,读出对齐可能是一个主导瓶颈。
英文摘要
Audio-conditioned language models often underuse acoustic cues such as prosody, emotion, and non-speech sounds, raising the question of whether ASR-supervised frontends discard this information before it reaches the LM. We test whether the frontend is responsible by comparing Whisper-Tiny and Whisper-Small with EnCodec, DAC-VAE, and WavTokenizer in a shared Qwen3.5-4B audio-LM pipeline on ASR, emotion recognition, and sound captioning. Encoder replacement alone does not resolve this underuse: Whisper variants remain strongest overall, including on emotion and environmental sound captioning. To localize the failure, we trace task-relevant information through the encoder, projector, LM layers, and LM head. Linear probes and geometric analyses show that discriminative acoustic structure remains recoverable at the final LM layer, even when MCQA accuracy trails probe accuracy by up to 83 points. Because the answer format and decoding procedure are controlled, this task-dependent gap points to content-specific readout failure rather than generic format bias. LogitLens analyses and a targeted LM head intervention support the conclusion that acoustic underuse is not explained solely by encoder-side information loss and that readout alignment can be a dominant bottleneck.