arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TAME:用于文本-视频检索的时间感知混合专家模型

TAME: Temporal-Aware Mixture-of-Experts for Text-Video Retrieval

Uicheol Jung, Juyoung Hong, Hojung Kwon, Yukyung Choi

arXiv 2609.02204首次发表:更新:

发表机构

Sejong University; Wisenut(世宗大学; 威森努特)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对文本-视频检索中时间建模缺失的问题,提出基于CLIP的TAME框架,通过MoE层、FT令牌和CTIA模块优化时间建模,在多个TVR基准上实现性能提升。

AI 中文摘要

文本-视频检索(TVR)旨在检索与自然语言查询匹配的视频,但将CLIP等图像-文本模型扩展至视频的核心限制在于缺乏时间建模能力。视频在外观和运动方面存在帧级异质性,将所有帧压缩为单一表示往往会掩盖时间结构与语义过渡。为解决该问题,我们提出TAME:用于文本-视频检索的时间感知混合专家模型(Temporal-Aware Mixture-of-Experts for Text-Video Retrieval,TAME),这是一种基于CLIP的框架,可联合建模帧级结构与时间关系。首先,我们在CLIP的两个编码器中集成稀疏混合专家(Mixture-of-Experts,MoE)层,并在视觉分支上应用帧一致路由,使专家根据帧级视觉模式进行专业化调整,同时保留原始的视觉-语言对齐。其次,我们引入帧-时间(Frame-Temporal,FT)令牌,用于聚合全局跨帧信息并将其反馈至每一帧,使视觉编码器能够捕捉长程时间依赖关系,且不会损害局部细节。第三,我们设计了跨时间交互与聚合(Cross-Temporal Interaction and Aggregation,CTIA)模块,通过分阶段时间过滤与融合来优化帧级句子-视频相似度。在标准TVR基准上的实验表明,TAME始终优于基于CLIP的基线模型:在MSR-VTT数据集上,其R@1指标较CLIP4Clip提升4.0;在DiDeMo、MSVD、LSMDC和ActivityNet数据集上也实现了一致的性能提升。代码可在该URL获取。

英文摘要

Text-Video Retrieval (TVR) retrieves videos that match a natural-language query, but extending image-text models such as CLIP to videos is fundamentally limited by the lack of temporal modeling. Videos exhibit frame-wise heterogeneity in appearance and motion, and compressing all frames into a single representation often obscures temporal structure and semantic transitions. To address this, we propose Temporal-Aware Mixture-of-Experts for Text-Video Retrieval (TAME), a CLIP-based framework that jointly models frame-level structure and temporal relations. First, we integrate sparse Mixture-of-Experts (MoE) layers into both CLIP encoders and apply frame-consistent routing on the vision branch so that experts specialize according to frame-level visual patterns while preserving the original vision-language alignment. Second, we introduce Frame-Temporal (FT) tokens that aggregate global cross-frame information and feed it back to each frame, enabling the visual encoder to capture long-range temporal dependencies without harming local details. Third, we design a Cross-Temporal Interaction and Aggregation (CTIA) module that refines frame-wise sentence-video similarities through staged temporal filtering and fusion. Experiments on standard TVR benchmarks show that TAME consistently improves over CLIP-based baselines. On MSR-VTT, it improves R@1 by 4.0 over CLIP4Clip, and also achieves consistent gains on DiDeMo, MSVD, LSMDC, and ActivityNet. The code is available at https://github.com/sejong-rcv/TAME.

Comments17 pages, 6 figures

Journal refIEEE Access 14 (2026) 16188-16203

DOI:10.1109/ACCESS.2026.3658103

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑