发表机构
KAIST(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对无训练视听事件感知的虚假共激活问题,提出SCoPE框架,通过跨模态标签竞争消除FCA,在LLP等数据集上显著提升性能且可直接迁移。
AI 中文摘要
视听事件感知(AVEP)用于确定视频中发生的事件、发生时间以及事件是否可听、可见或两者兼具。无训练方法通过将冻结的音频和视觉特征与文本编码的事件名称匹配来查询新的事件词汇,但相关标签会共享证据,导致错误标签的得分至少与正确标签一样高,我们将此称为虚假共激活(FCA)。没有标量阈值能在保留所有正确标签的同时拒绝错误标签,特定类别的阈值可能阻止该标签成为最终预测,但FCA仍存在于底层得分向量中。我们提出SCoPE,这是一个无训练框架,其中所有查询标签竞争共享证据,且每个模态引导另一个模态的事件选择。我们推导了两标签拟合中此竞争消除FCA的精确条件。在LLP上使用相同的冻结CLIP+CLAP骨干网络时,与已报道的AV²A值相比,SCoPE的Type@seg指标提升7.45个百分点,Event@seg指标提升5.04个百分点,相同的固定配置可直接迁移至OV-AVEBench和VGGSound-AVEL100k数据集。
英文摘要
Audio-visual event perception (AVEP) determines which events occur in a video, when they occur, and whether they are audible, visible, or both. Training-free methods query new event vocabularies by matching frozen audio and visual features with text-encoded event names. However, related labels share evidence. An incorrect label can then score at least as high as a correct one. We call this a false co-activation (FCA). No scalar cutoff can reject the incorrect label while keeping every correct one. Class-specific thresholds may prevent that label from becoming a final prediction, but the FCA remains in the underlying score vector. We introduce SCoPE, a training-free framework in which all queried labels compete for shared evidence and each modality guides event selection in the other. We derive an exact condition for when this competition removes an FCA in a two-label fit. With identical frozen CLIP+CLAP backbones on LLP, SCoPE improves Type@seg by 7.45 points and Event@seg by 5.04 points compared with the reported AV$^2$A values. The same fixed configuration transfers unchanged to OV-AVEBench and VGGSound-AVEL100k.
Comments17 pages, 6 figures