arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.01778cs.CV

先分配再嵌入:面向视频嵌入的自适应视觉输入分配

Allocate Before You Embed: Adaptive Visual Input Allocation for Video Embeddings

Song Jin, Zhongtao Jiang, Chenglei Shen, Huanxuan Liao, Haozhe Chi, Zhiwei Wang, Kun Xu, Yong Liu

首次发表
浏览论文内容

中文总结 AI 辅助

针对大规模视频检索的视觉输入预算限制问题,提出AllocEmbed框架,通过RDPO学习分配器重分配预算,在多基准任务上实现预算匹配方法中最佳检索性能且可跨主干迁移。

中文摘要 AI 辅助

大规模视频检索要求嵌入模型在严格的视觉输入和推理预算下对长且多样的视频进行编码。现有方法通常以原始分辨率采样少量固定数量的帧,这限制了时间覆盖范围且忽略了帧的重要性。我们的实证分析表明,即使在固定视觉输入预算下,扩大时间覆盖范围也能提升检索性能;若保留原始单帧分辨率,提升幅度会更大,凸显了时间覆盖与空间保真度的互补作用。受该发现启发,我们提出了AllocEmbed,一种先分配再嵌入的框架,它会在更多帧之间重新分配固定的视觉输入预算。轻量型分配器利用低成本预览在嵌入主干前分配逐帧分辨率,在对检索最有益的区域保留更多细节,同时在其他区域降低视觉成本。我们还提出了检索驱动策略优化(RDPO),它利用经排序验证的相似度差距和置信度引导的效率激励,直接从检索反馈中学习分配器。AllocEmbed完全在主干前运行,可与现有检索系统集成,无需修改嵌入模型或下游流程。在MMEB-V2 V-QA、V-RET任务及我们的LongRet基准上的实验表明,AllocEmbed在所有评估的预算匹配方法中实现了最佳整体检索性能,且可跨嵌入主干迁移。我们的代码可在该https URL获取。

英文摘要

Large-scale video retrieval requires embedding models to encode long and diverse videos under tight visual-input and inference budgets. Existing methods typically sample a small, fixed set of frames at their original resolution, limiting temporal coverage and ignoring frame importance. Our empirical analysis shows that expanding temporal coverage improves retrieval even under a fixed visual-input budget. Gains are larger when the original per-frame resolution is preserved, highlighting the complementary roles of temporal coverage and spatial fidelity. Motivated by this finding, we propose AllocEmbed, an allocate-then-embed framework that reallocates a fixed visual-input budget across more frames. A lightweight allocator uses low-cost previews to assign frame-wise resolutions before the embedding backbone, preserving more detail where it most benefits retrieval while reducing visual cost elsewhere. We further introduce Retrieval-Driven Policy Optimization (RDPO), which learns the allocator directly from retrieval feedback using a rank-validated similarity gap and a confidence-guided efficiency incentive. Operating entirely before the backbone, AllocEmbed integrates with existing retrieval systems without modifying the embedding model or downstream pipeline. Experiments on the MMEB-V2 V-QA and V-RET tasks and our LongRet benchmark show that AllocEmbed achieves the best overall retrieval performance among the evaluated budget-matched methods and transfers across embedding backbones. Our code is publicly available at https://github.com/jinsong8/AllocEmbed.

发表机构

  • Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学高瓴人工智能学院)
  • Institute of Automation, Chinese Academy of Science(中国科学院自动化研究所)
  • Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

↑