arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OP-CAD:用于稳健视听推理的在线策略干净音频蒸馏

OP-CAD: On-Policy Clean-Audio Distillation for Robust Audio-Visual Reasoning

Xingming Shui, Dapeng Chen, Bowei Liu, Jingqi Tian, Minfu Li, Kun Yi, Jiapeng Hong, Yansong Tang

arXiv 2609.39150首次发表:更新:

发表机构

Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对视听问答中的噪声干扰问题,提出基于课程学习的在线策略干净音频蒸馏框架OP-CAD,通过选择性token级监督提升稳健性,优于现有方法且不损失干净准确率。

AI 中文摘要

部署在真实环境中的全模态大语言模型会遇到外部噪声,这些噪声可能干扰其对多模态输入的感知和理解。我们研究了它们在视听理解中的稳健性,重点关注环境噪声和竞争语音下的问答任务。挑战在于抵抗声学干扰的同时保留有用的音频证据。在线策略蒸馏为学生生成的响应提供了密集的教师反馈,但均匀的token加权并未明确优先考虑受声学干扰影响的位置。我们提出了OP-CAD(在线策略干净音频蒸馏),一种基于课程学习的特权自蒸馏框架,用于稳健的视听理解。训练从轻度到重度环境噪声和竞争语音逐步推进,每个阶段都进行选择性的token级监督。学生从受损的视听输入生成响应,而冻结的教师使用干净音频和验证答案来监督相同的响应前缀。为了分配这种监督,OP-CAD在不泄露答案的情况下比较教师在干净、受损和仅视觉上下文下的预测。这些匹配比较衡量了对音频移除和损坏的敏感性;一个有界的加权规则强调由任一信号识别出的位置,同时在整个响应过程中保留监督。OP-CAD在所有评估的噪声条件下均优于比较方法。配对分析进一步表明,在强干扰下干净正确答案的保留得到改善,且未观察到总体干净准确率的损失。这些结果证明了将干净教师监督导向声学敏感预测对于稳健视听推理的价值。

英文摘要

Omni-modal large language models deployed in real-world environments encounter external noise that can interfere with their perception and understanding of multimodal inputs. We study their robustness in audio-visual understanding, focusing on question answering under environmental noise and competing speech. The challenge is to resist acoustic interference while preserving useful audio evidence. On-policy distillation provides dense teacher feedback on student-generated responses, but uniform token weighting does not explicitly prioritize positions affected by acoustic interference. We introduce OP-CAD (On-Policy Clean-Audio Distillation), a curriculum-based privileged self-distillation framework for robust audio-visual understanding. Training progresses from mild to severe environmental noise and competing speech, with selective token-level supervision at each stage. The student generates responses from corrupted audio-visual input, while a frozen teacher uses clean audio and the verified answer to supervise the same response prefixes. To allocate this supervision, OP-CAD compares teacher predictions under clean, corrupted, and visual-only contexts without revealing the answer. These matched comparisons measure sensitivity to audio removal and corruption; a bounded weighting rule emphasizes positions identified by either signal while retaining supervision throughout the response. OP-CAD outperforms the compared methods across all evaluated noise conditions. Paired analyses further show improved preservation of clean-correct answers under strong interference, with no observed aggregate clean-accuracy penalty. These results demonstrate the value of directing clean-teacher supervision toward acoustically sensitive predictions for robust audio-visual reasoning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑