arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

增强长视频VLM嵌入的查询感知流式潜在推理

Enhancing Long-Video VLM Embeddings with Query-Aware Streaming Latent Reasoning

Haozhe Chi, Song Jin, Yang Jin, Yadong Mu

arXiv 2610.04864首次发表:更新:

发表机构

Peking University; Renmin University of China(北京大学; 中国人民大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出查询感知流式潜在推理(QASLR),通过跨片段累积证据并保持嵌入大小固定,显著提升长视频检索性能,HourVideo Hit@1最高提升至72.8。

AI 中文摘要

长视频嵌入需要在有限的视觉令牌预算下捕获稀疏的查询相关证据。均匀采样可能会错过跨越数分钟或数小时的视频中的短暂事件,而在单个上下文中编码更多帧会增加内存和计算量。我们提出了查询感知流式潜在推理(QASLR),这是一种后训练框架,在保持嵌入大小固定的同时,跨片段累积证据。QASLR选择一组有界的帧,恢复其时间顺序,并使用视觉-语言骨干逐片段处理它们。一组紧凑的持久思考令牌与每个片段的特征进行交叉注意力,而嵌入令牌在每次更新后读出归一化表示。这种设计在多次骨干调用中整合证据,而无需所有选定帧共享单个上下文。训练结合了最终对比学习、逐步对比监督和最终嵌入自蒸馏。中间监督训练部分视频读出以用于检索,而自蒸馏将其正则化以朝向最终表示。查询感知选择产生查询条件表示,用于候选集评分和重排序,而查询无关选择则支持可重用的语料库索引。在完整训练方案下,HourVideo检索的Hit@1分别从54.2提高到70.7,从57.4提高到72.8,对于2B和8B Qwen3-VL-Embedding骨干。增益扩展到评估的时刻检索和视频问答任务,流式头转移到第二个Qwen系列嵌入骨干。这些结果支持流式潜在聚合作为将长视频证据整合到固定维度表示中的有效方法。

英文摘要

Long-video embedding requires capturing sparse query-relevant evidence under a limited visual-token budget. Uniform sampling can miss brief events in videos spanning minutes or hours, whereas encoding more frames in a single context increases memory and computation. We introduce \textbf{Query-Aware Streaming Latent Reasoning} (QASLR), a post-training framework that accumulates evidence across clips while keeping the embedding size fixed. QASLR selects a bounded set of frames, restores their temporal order, and processes them clip by clip with a vision-language backbone. A compact set of persistent think tokens cross-attends to each clip's features, while an embed token reads out a normalized representation after every update. This design integrates evidence across multiple backbone calls without requiring all selected frames to share a single context. Training combines final contrastive learning, step-wise contrastive supervision, and final-embedding self-distillation. Intermediate supervision trains partial-video readouts for retrieval, while self-distillation regularizes them toward the final representation. Query-aware selection produces query-conditioned representations for candidate-set scoring and reranking, whereas query-independent selection enables reusable corpus indexing. Under the full training recipe, HourVideo retrieval Hit@1 increases from 54.2 to 70.7 and from 57.4 to 72.8 for 2B and 8B Qwen3-VL-Embedding backbones, respectively. Gains extend to the evaluated moment-retrieval and video-QA tasks, and the streaming head transfers to a second Qwen-family embedding backbone. These results support streaming latent aggregation as an effective approach to integrating long-video evidence into fixed-dimensional representations.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑