大型音频语言模型中用于声学感知的编码器端神经元识别与增强
Encoder-Side Neuron Identification and Amplification for Acoustic Perception in Large Audio-Language Models
浏览论文内容
中文总结 AI 辅助
研究针对大型音频语言模型在语音非语义属性表现不佳的问题,提出IAAN方法,通过对比音频编码器神经元在真实波形与噪声参考上的激活来评分并增强高分神经元,有效提升了模型在多数据集上的声学感知准确率,开辟了新的推理时改进方向。
中文摘要 AI 辅助
大型音频语言模型在语音内容方面表现出色,但在语音细粒度、非语义属性(如说话者情感)上往往表现不佳。在不重新训练的情况下改进这一点需要有效的推理时干预,然而大多数现有方法仅在音频编码器之后进行干预,且粒度较粗。编码器本身在很大程度上未被探索,特别是在单个神经元层面。我们引入了IAAN(识别和增强声学神经元),一种无需训练和标签的方法,通过将音频编码器中每个前馈神经元在真实波形上的激活与缺乏真实音频声学信息的噪声参考上的激活进行对比来评分。然后在推理时增强一小部分得分最高的神经元。在十个非语义语音属性上,IAAN在Audio - Flamingo - 3上平均准确率提高25.7分,在Qwen2.5 - Omni上提高21.4分,在Kimi - Audio上提高9.7分。它还能提升已明确微调以优先考虑声学证据的模型。在控制比较中,编码器位置和神经元级选择性对这种提升都很必要。在编码器之后、解码端或语言模型内部进行干预几乎没有改善甚至会降低准确率。改进还取决于增强哪些特定神经元,而非仅仅数量。这些结果表明,在音频编码器内部进行小而精准的干预是加强大型音频语言模型声学理解的有效且未充分利用的方法,为通过神经元级访问编码器来改善声学感知的推理时方法开辟了新方向。
英文摘要
Large audio-language models (LALMs) often underperform on fine-grained, non-semantic attributes of speech, such as a speaker's emotion, despite strong performance on speech content. Improving this without the cost of retraining calls for an effective inference-time intervention, yet most existing methods intervene only after the audio encoder and operate at a relatively coarse granularity. The encoder itself, where acoustic information is first extracted from the waveform, remains largely unexplored, especially at the level of individual neurons. We introduce IAAN, Identifying and Amplifying Acoustic Neurons, a training-free and label-free method that scores each feed-forward neuron in the audio encoder by contrasting its activation on the real waveform with that on a noise reference lacking the real audio's acoustic information. IAAN then amplifies a small set of the highest-scoring neurons at inference. Across ten non-semantic speech attributes, IAAN improves average accuracy by 25.7 points on Audio-Flamingo-3, 21.4 on Qwen2.5-Omni, and 9.7 on Kimi-Audio. It also improves a model already explicitly fine-tuned to prioritize acoustic evidence. In controlled comparisons, both the encoder locus and neuron-level selectivity prove necessary for this gain. Intervening after the encoder, at the decoding side or inside the language model, yields little to no improvement, or even deteriorates accuracy. The improvement also depends on which specific neurons are amplified, not merely on their number, confirming that IAAN's acoustic score succeeds in identifying the neurons that matter. These results show that a small, precisely targeted intervention inside the audio encoder is an effective and largely untapped way to strengthen the acoustic understanding of LALMs, opening a new direction for inference-time methods that improve acoustic perception through neuron-level access to the encoder.
发表机构
- National Taiwan University(国立台湾大学)
- Graduate Institute of Communication Engineering(通信工程研究所)
- NTU Artificial Intelligence Center of Research Excellence(国立交通大学人工智能卓越研究中心)
- ASUS Open Cloud Infrastructure Software Center(ASUS开放云基础设施软件中心)
机构由 AI 辅助整理,请以论文原文为准。