发表机构
IIT Kanpur; Dolby Laboratories(坎普尔印度理工学院; 杜比实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对开放词汇视听事件定位任务,提出基于复值神经网络的四流融合架构,在OV-AVEBench和修改后的AVE数据集上均实现最优性能,双流方案也有显著提升。
AI 中文摘要
开放词汇视听事件定位(OV-AVEL)为每个视频片段标注事件类别,包含训练期间从未见过的类别。主流流程采用冻结的多模态基础模型(如ImageBind)将视觉帧、音频梅尔频谱图及每个候选类别名称嵌入共享空间,随后计算每个片段与每个类别的两个余弦相似度:视觉-文本相似度和音频-文本相似度。现有方法会通过固定规则(几何均值、加权平均)将该对值合并为单个标量分数,再取最大值。相反,我们计算复值相似度并使用复值神经网络(CVNN)学习其融合。每个模态的标准表示成为我们流程的实部,配对的辅助流提供虚部;我们使用iHSV的虚部作为视觉模态的辅助流,使用CycleGAN转换的相位频谱图作为音频模态的辅助流,由此得到两个复值相似度并进行融合。视觉和音频编码器保持冻结,仅训练时间注意力块和融合CVNN。该四流复值架构在两个OV-AVEL基准上达到新的最优性能:在OV-AVEBench的开放(未见过类别)划分上,我们的方法达到66.5%准确率、59.1%分段F1值、54.1%事件F1值,较之前报告的微调基线分别提升1.6、4.1、6.6个百分点,对见过类别也有一致提升。我们还为该任务修改了AVE数据集,观察到我们的架构达到60.7%准确率、51.9%分段F1值、50.4%事件F1值,在该数据集上也实现了最优的OV-AVEL结果;我们还提出了一种双流替代方案,其性能也较基线有大幅提升。
英文摘要
Open-Vocabulary Audio-Visual Event Localization (OV-AVEL) labels each video segment with an event class, including classes that were never seen during training. The dominant pipeline uses a frozen multimodal foundation model (e.g. ImageBind) to embed the visual frame, the audio mel-spectrogram, and each candidate class name into a shared space, then computes two cosine similarities for each segment against each class: visual-text and audio-text. Existing methods then collapse this pair into a single scalar score with a fixed rule (geometric mean, weighted average) before taking the argmax. Instead, we compute complex-valued similarities and learn their fusion using a complex-valued neural network (CVNN). Each modality's standard representation becomes the real part of our pipeline, and a paired companion stream supplies the imaginary part. We use imaginary part of iHSV for visual modality and CycleGAN-translated phase spectrogram for audio modality as these companion streams. This results in two complex similarities, which are then fused. While the vision and audio encoders remain frozen, only the temporal-attention blocks and the fusion CVNN are trained. The four-stream complex architecture sets a new state of the art on both OV-AVEL benchmarks. On the open (unseen-class) split of OV-AVEBench we reach 66.5/59.1/54.1% Acc/Seg-F1/Event-F1 (+1.6/+4.1/+6.6 over the previously reported fine-tuned baseline), with consistent gains for seen classes as well. We also modify AVE dataset for this task and observe that our architecture reaches 60.7/51.9/50.4% Acc/Seg-F1/Event-F1, achieving state-of-the-art OV-AVEL results on it as well. We also propose a two-stream alternative, which also sees great improvements over the baseline.
CommentsAccepted to British Machine Vision Conference (BMVC) 2026