发表机构
Concordia University; Mila – Quebec AI Institute; Technion – IIT; Laval University(康考迪亚大学; 米拉-魁北克人工智能研究所; 以色列理工学院; 拉瓦尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AnchorPrompt通过冻结模型并学习解码器输入处的软提示向量,利用自蒸馏训练,在多种扰动下提升音频-语言模型的答案一致性并减少幻觉,且支持零样本迁移到未见扰动。
AI 中文摘要
大型音频-语言模型(LALMs)对输入扰动(如噪声、波形损坏和对抗性注入)敏感。我们提出AnchorPrompt,一种高效的适配方法,该方法保持模型冻结,并在解码器输入处、音频嵌入和问题嵌入之间学习单个提示向量块。我们通过在不同音频和文本扰动上的自蒸馏来训练这些向量。为提高答案一致性并缓解幻觉,我们使用模型对干净录音的预测作为可回答输入的目标,并在音频缺乏足够证据回答时分配拒绝目标。此外,AnchorPrompt在推理时对扰动不可知,无需事先检测扰动,并能零样本迁移到未见过的失真。我们在三个基准上评估了三个LALM,并显示AnchorPrompt在大多数测试条件下提高了答案一致性。在九个模型-基准对中,六个的干净准确率得到提升,其余部分的影响最多为1.2%。关键的是,AnchorPrompt在严重音频损坏下减少幻觉,同时保持干净音频上的错误拒绝罕见。最后,这些一致性收益可迁移到未见过的扰动,如选择排列和混响。
英文摘要
Large audio-language models (LALMs) are sensitive to input perturbations, such as noise, waveform corruption, and adversarial injections. We propose AnchorPrompt, an efficient adaptation method that keeps the model frozen and learns a single block of prompt vectors inserted at the decoder input, between the audio and question embeddings. We train these vectors through self-distillation over diverse audio and text perturbations. To improve answer consistency and mitigate hallucination, we use the model's prediction on the clean recording as the target for answerable inputs, and assign a refusal target when the audio lacks sufficient evidence to answer. Furthermore, AnchorPrompt is perturbation-agnostic at inference, requiring no prior detection of perturbations and enabling zero-shot transfer to unseen distortions. We evaluate three LALMs across three benchmarks and show that AnchorPrompt improves answer consistency in most tested conditions. Clean accuracy improves in six of nine model-benchmark pairs, with minimal impact on the remainder of 1.2% at most. Crucially, AnchorPrompt reduces hallucinations under severe audio corruption while keeping false refusals on clean audio rare. Finally, these consistency gains transfer to unseen perturbations, such as choice permutations and reverberation.