CurvSpec:面向部分相关视频检索的自适应多曲率学习
CurvSpec: Adaptive Multi-Curvature Learning for Partial Relevant Video Retrieval
浏览论文内容
中文总结 AI 辅助
CurvSpec提出内容自适应曲率学习框架,通过并行欧几里得与双曲注意力及语义质心匹配,解决部分相关视频检索中的信号稀释与曲率刚性挑战,在三个基准上达到最优性能。
中文摘要 AI 辅助
部分相关视频检索(PRVR)旨在检索包含与文本查询匹配的瞬间的未修剪视频,且无需时间标注。相关瞬间可能仅持续几秒,而视频长达数分钟,导致信噪比极低,使得PRVR比标准的全视频检索更具挑战性。该任务面临两个相互交织的挑战:(1)信号稀释,即粗糙的全局表示将短暂的相关信号淹没在占主导地位的不相关背景中;(2)曲率刚性,即将所有视频嵌入同一固定几何空间会扭曲从平坦原子事件到深层组合层次结构的各类视频的表示。现有PRVR方法改进了瞬间选择和跨模态匹配,但通常仍将所有视频编码在单一固定曲率的检索空间中,限制了其建模多样化视频结构的能力。为解决这两个挑战,我们提出CurvSpec,一个学习内容自适应曲率用于视频检索表示的框架,而非强加固定几何先验。CurvSpec通过并行的欧几里得和双曲注意力层处理特征,为双曲层分配独立学习的曲率,并通过内容感知融合机制将每个输入路由到最合适的几何区域。为进一步抑制信号稀释,CurvSpec用语义质心表示每个视频,质心数量由视频内容复杂度决定,将其投影到学习到的流形上,并通过测地距离将每个查询与最近的质心匹配。在ActivityNet Captions、TVR和Charades-STA上的实验展示了最先进的检索性能。
英文摘要
Partially Relevant Video Retrieval (PRVR) seeks to retrieve untrim-med videos containing a moment that matches a text query, without temporal annotations. The relevant moment may last only seconds within a video spanning several minutes, creating an extremely low signal-to-noise ratio that makes PRVR more challenging than standard full-video retrieval. This task presents two intertwined challenges: (1) signal dilution, where coarse global representations blur the brief relevant signal into the dominant irrelevant surroundings;(2) curvature rigidity, where embedding all videos in the same fixed-geometry space distorts representations for videos that range from flat atomic events to deep compositional hierarchies. Existing PRVR methods have improved moment selection and cross-modal matching, but they still typically encode all videos in a single fixed-curvature retrieval space, limiting their ability to model diverse video structures. To address both challenges, we propose CurvSpec, a framework that learns content-adaptive curvature for video retrieval representations rather than imposing a fixed geometric prior. CurvSpec processes features through parallel Euclidean and hyperbolic attention layers, with independently learned curvatures assigned to the hyperbolic layers, and a content-aware fusion mechanism routes each input to its most suitable geometric regime. To further suppress signal dilution, CurvSpec represents each video with semantic centroids whose number is determined by the video's content complexity, projects them onto the learned manifold, and matches each query against its nearest centroid by geodesic distance. Experiments on ActivityNet Captions, TVR, and Charades-STA demonstrate state-of-the-art retrieval performance.
发表机构
- Tongji University(同济大学)
- SIGS, Tsinghua University(清华大学深圳国际研究生院)
- Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。