AI 中文总结
FATE是帧级视听时间嵌入模型,通过保留帧级序列与联合对比学习目标,在三项视听任务中优于基线,零样本事件定位性能接近全监督方法,与人类判断相关性最优。
AI 中文摘要
当狗张嘴吠叫时,人类能自然识别声音内容与发生时间,构建具备该能力的视听模型需要同时捕捉语义与时间对齐的表征。现有方法存在不足:嵌入模型能匹配语义但丢失时间信息;同步模型能捕捉时间偏移但缺乏语义理解。为弥合差距,我们提出FATE(Frame-level Audio-visual Temporal Embedding,帧级视听时间嵌入)。与将各模态池化为单个嵌入并丢弃时间信息的现有嵌入模型不同,FATE保留帧级序列,在物理时间线上对齐,并对严格对齐的帧对计算相似度。与仅输出偏移预测的同步模型不同,FATE在可复用嵌入空间中编码同步信息,通过结合跨视频语义对比学习与视频内时间对比学习的联合目标训练,以同时捕捉声音内容与发生时间。在三项任务中,FATE在时间与语义检索上大幅超越最强基线,在零样本设置下的事件定位上匹配全监督方法的性能,且作为生成评估指标时与人类判断的相关性最优。源代码可在该网址获取:https://this.url
英文摘要
When a dog opens its mouth and barks, humans naturally recognize what the sound is and when it occurs. Building audio-visual models with this same ability requires representations that capture both semantic and temporal alignment. Current approaches fall short on one side or the other: embedding models match semantic but lose temporal information; synchronization models capture temporal offsets but lack semantic understanding. To bridge this gap, we propose FATE, Frame-level Audio-visual Temporal Embedding. Unlike prior embedding models that pool each modality into a single embedding and discard temporal information, FATE retains frame-level sequences, aligns them on the physical timeline, and computes similarity over strictly aligned frame pairs. Unlike synchronization models that output only an offset prediction, FATE encodes synchronization in a reusable embedding space, trained with a joint objective combining cross-video semantic and within-video temporal contrastive learning to capture both what sounds and when it occurs. Across three tasks, FATE surpasses the strongest baseline on temporal and semantic retrieval by a large margin, matches fully supervised methods on event localization in a zero-shot setting, and achieves the best correlation with human judgments as a generation evaluation metric. The source code can be found at \texttt{https://github.com/guankaisi/FATE}.