arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

如果你听到它,就帮助找到它:面向开放词汇音频-视觉事件定位的正交知识蒸馏

If You Hear It, Help Find It: Orthogonal Knowledge Distillation for Open-Vocabulary Audio-Visual Event Localization

Yi Xu, Cheng Chen, Wenzhuo Lei

arXiv 2609.23376首次发表:更新:

发表机构

Tongji University(同济大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对开放词汇音频-视觉事件定位中教师监督可靠性差异问题,提出可靠性感知的非对称蒸馏框架OV-OrthKD,通过正交损失分离视觉与音频教师投影,在OV-AVEBench上显著提升定位性能。

AI 中文摘要

开放词汇音频-视觉事件定位(OV-AVEL)旨在从视频、音频和语言中定位文本查询事件的时间边界。该任务可用的监督源在时间边界可靠性上存在差异:在OV-AVEBench上,我们配置的视觉教师模型比配置的音频教师模型提供更可靠的边界线索,尽管后者是一个强大的预训练音频模型且在语义上仍具信息量。这是特定设置下的诊断结果,而非视觉与音频的通用排序。我们将由此产生的挑战表述为监督放置问题:哪些教师信号应塑造定位决策,哪些应保持辅助性。基于此观点,我们提出OV-OrthKD,一种可靠性感知的非对称蒸馏框架。视觉特征迁移塑造决策对齐的表示,音频特征迁移丰富互补的辅助子空间,文本原型锚定已见/未见类别语义,正交损失限制两个教师特定投影之间的方向重叠。学生模型在推理时通过查询感知融合继续使用两种模态,而默认训练方案使音频教师监督不进入片段逻辑路径。在OV-AVEBench上,OV-OrthKD达到0.816的片段AP,并在F1@0.5上相比官方微调基线整体提升2.7个百分点,未见类别提升3.4个百分点。路径分配、角色交换、损坏和迁移分析一致支持监督放置作为OV-AVEL的任务特定设计轴。

英文摘要

Open-vocabulary audio-visual event localization (OV-AVEL) grounds a text-queried event in time from video, audio, and language. The supervision sources available to this task can differ in temporal-boundary reliability: on OV-AVEBench, our configured visual teacher gives more reliable boundary cues than the configured audio teacher, although the latter is a strong pretrained audio model and remains semantically informative. This is a setting-specific diagnostic rather than a universal ranking of vision and audio. We formulate the resulting challenge as supervision placement: which teacher signals may shape the localization decision, and which should remain auxiliary. Based on this view, we propose OV-OrthKD, a reliability-aware asymmetric distillation framework. Visual feature transfer shapes a decision-aligned representation, audio feature transfer enriches a complementary auxiliary subspace, a text prototype anchors seen/unseen category semantics, and an orthogonality loss limits directional overlap between the two teacher-specific projections. The student continues to use both modalities through query-aware fusion at inference, while the default training recipe keeps audio-teacher supervision off the segment-logit path. On OV-AVEBench, OV-OrthKD achieves 0.816 segment AP and improves F1@0.5 over the official fine-tuning baseline by 2.7 points overall and 3.4 points on unseen categories. Path-assignment, role-swap, corruption, and transfer analyses consistently support supervision placement as a task-specific design axis for OV-AVEL.

CommentsAccepted to ACM Multimedia 2026 (poster). 9 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑