重新思考长视频效率:帧、像素与前端口延迟的联合分配视角
Rethinking Long-Video Efficiency: A Joint Allocation Perspective on Frames, Pixels, and Front-End Latency
浏览论文内容
中文总结 AI 辅助
提出LoHi框架,通过密集低分辨率视频流与稀疏高分辨率图像帧的联合分配,在匹配令牌预算下提升长视频理解准确率,并大幅降低前端口延迟。
中文摘要 AI 辅助
高效的长视频理解与视觉语言模型(VLMs)通常被表述为在固定原生分辨率下选择信息丰富的帧或视觉令牌。我们表明,每帧分辨率可以转而用于换取更密集的时间覆盖,而前端口解码延迟取决于候选池的大小,而非最终的令牌预算。在多个VLMs和长视频基准上的实证研究得出了三个发现:在匹配的令牌预算下,密集低分辨率采样优于稀疏原生分辨率采样;对分辨率敏感的任务受益于选定的高分辨率帧;前端口解码主导了小时级视频的墙钟时间。受这些发现启发,我们提出了LoHi,一个无需训练、单遍的框架,通过VLM的原生视频和图像路径,将密集低分辨率视频流与稀疏高分辨率图像帧相结合。LoHi-Anchor使用编解码器I帧元数据选择高分辨率帧,而LoHi-SemDiv使用查询相关性和CLIP特征上的视觉多样性。在三个长视频基准上,LoHi在匹配的令牌预算下,相比原生分辨率基线平均准确率提高了10.6个百分点,相比最强先验效率方法提高了5.2个百分点。它还将小时级视频的前端口解码延迟降低了最多7倍。项目页面:此https URL
英文摘要
Efficient long-video understanding with vision-language models (VLMs) is often framed as selecting informative frames or visual tokens at a fixed native resolution. We show that per-frame resolution can instead be traded for denser temporal coverage, while front-end decoding latency depends on the size of the candidate pool rather than the final token budget. An empirical study across multiple VLMs and long-video benchmarks yields three findings: dense low-resolution sampling outperforms sparse native-resolution sampling at matched token budgets; resolution-sensitive tasks benefit from selected high-resolution frames; and front-end decoding dominates wall time for hour-long videos. Motivated by these findings, we introduce LoHi, a training-free, single-pass framework that combines a dense low-resolution video stream with sparse high-resolution image frames through the VLM's native video and image pathways. LoHi-Anchor selects high-resolution frames using codec I-frame metadata, while LoHi-SemDiv uses query relevance and visual diversity over CLIP features. Across three long-video benchmarks, LoHi improves average accuracy by 10.6 percentage points over the native-resolution baseline at a matched token budget and by 5.2 percentage points over the strongest prior efficiency method. It also reduces front-end decoding latency by up to 7x on hour-long videos. Project page: https://sixundong.com/projects/lohi
发表机构
- University of Central Florida(中佛罗里达大学)
- Meta Reality Labs(Meta现实实验室)
- Axon(Axon公司)
机构由 AI 辅助整理,请以论文原文为准。