音频语言模型中基于长期记忆引导的目标感知增强
Long-Term Memory-Guided Enhancement for Target Perception in Audio-Language Models
浏览论文内容
中文总结 AI 辅助
提出LTM-AE,利用人类长期记忆原理,无需训练即可增强音频大语言模型在噪声环境中的目标感知,显著提升分类准确率并降低语音识别错误率。
中文摘要 AI 辅助
音频大语言模型(ALLMs)能够对音频录制内容进行推理以执行复杂任务。然而,在现实环境中,当背景噪声和竞争声源混合目标声音时,这些能力通常会失效。受人类听觉中长期记忆的启发,我们提出了长期记忆引导的音频增强(LTM-AE),通过无需训练地优化ALLMs的音频表示来改善选择性目标感知。LTM-AE从每个类别的独立干净参考录音中提取隐藏状态中的表示作为长期记忆,引导增强朝向用户指定的聆听目标。我们在选定类别的长期记忆中重建传入的音频标记,并在语言主干解码之前将重建结果与原始标记进行插值。这种插值控制存储的听觉经验的影响,同时保持所有ALLM参数固定。在二十个声音类别和三个ALLMs上的诊断读数显示,LTM-AE在三种干扰源中增强了对指定目标的响应。在约束和自由形式分类的平均情况下,多个开源模型在原始混合上的准确率提升范围从29.53到46.15个百分点。对于语音内容恢复,带有额外学习令牌级门控的LTM-AE将Qwen2-Audio的词错误率从23.07%降低到14.77%。这项工作朝着利用人类长期记忆原理增强ALLMs以实现现实世界聆听迈出了初步一步。我们的代码可在以下https URL获取。
英文摘要
Audio large language models (ALLMs) can reason about the content of audio recordings to perform complex tasks. However, these capabilities usually collapse in real-world environments when background noise and competing sources mix the target sound. Inspired by long-term memory in human listening, we propose Long-Term Memory-Guided Audio Enhancement (LTM-AE) to improve selective target perception by refining the audio representations of ALLMs without training. LTM-AE extracts representations in hidden states from separate clean reference recordings as long-term memory for each category, guiding enhancement toward a user-specified listening target. We reconstruct incoming audio tokens in the selected category long-term memory and interpolate the reconstructions with the original tokens before language backbone decoding. This interpolation controls the influence of stored auditory experience while keeping all ALLM parameters fixed. Diagnostic readouts across twenty sound categories and three ALLMs show that LTM-AE strengthens responses to a specified target amid three interfering sources. Averaged over constrained and free-form classification, accuracy gains over raw mixtures range from 29.53 to 46.15 percentage points across multiple open source models. For speech content recovery, LTM-AE with an additional learned token-level gate reduces Qwen2-Audio's word error rate from 23.07% to 14.77%. This work takes an initial step toward using principles of human long-term memory to enhance ALLMs for real-world listening. Our code is available at https://github.com/aynlp/ltm-audio-code
发表机构
- Nanyang Technological University(南洋理工大学)
- Beijing University of Posts and Telecommunications(北京邮电大学)
机构由 AI 辅助整理,请以论文原文为准。