arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.00247cs.SDcs.AI

对比音频解码的自适应扰动选择

Adaptive Perturbation Selection for Contrastive Audio Decoding

Aaron Isidore Grace, Zhouyuan Huo, Weiran Wang

首次发表
浏览论文内容

中文总结 AI 辅助

针对大型音频语言模型幻觉问题,提出自适应选择最优音频扰动作为对比解码负分支的方法,在时序、存在性等任务上提升准确率。

中文摘要 AI 辅助

大型音频语言模型(LALMs)经常通过用语言先验覆盖声学证据而产生幻觉。虽然对比解码(CD)提供了无需训练的缓解方法,但现有方法依赖于掩蔽或噪声等粗糙扰动,忽略了结构化音频变换。我们通过评估多样化的目标音频扰动库,并为每个任务和示例自适应选择最优负分支来探索这一设计空间。首先,我们改进了早期的提示工程,表明简单的二元是/否约束减少了模型错误确认缺失音频特征的倾向。其次,在时域、谱域、频域和幅度域上评估我们的库,发现最优变换高度依赖于任务;例如,反转音频数组破坏了时间连贯性,将时序任务的准确率从74.7%提高到81.4%。最后,我们在模型隐藏状态上训练了一个轻量级扰动选择器,以动态路由负分支,在存在性任务上额外获得了+4.3%的提升。

英文摘要

Large audio-language models (LALMs) frequently hallucinate by overriding acoustic evidence with language priors. While contrastive decoding (CD) offers training-free mitigation, existing methods rely on blunt perturbations like masking or noise, leaving structured audio transformations unexplored. We explore this design space by evaluating a diverse library of targeted audio perturbations and adaptively selecting the optimal negative branch for each task and example. First, we improve upon earlier prompt engineering by showing that a simple binary yes/no constraint reduces the model's tendency to falsely confirm absent audio features. Second, evaluating our library across temporal, spectral, frequency, and amplitude domains reveals that optimal transformations are highly task-dependent; for instance, reversing the audio array disrupts temporal coherence, raising accuracy on the temporal order task from 74.7% to 81.4%. Finally, we trained a light-weight perturbation selector on model hidden states to dynamically route negative branches, yielding an additional +4.3% gain on the existence task.

发表机构

  • Google(谷歌)
  • University of Iowa(爱荷华大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑