arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LENS:用于长视频关键帧采样的自适应时空缩放

LENS: Adaptive Spatio-Temporal Zooming for Keyframe Sampling in Long-Form Videos

Ce Zhang, Jinxi He, Katia Sycara, Yaqi Xie

arXiv 2607.25125首次发表:更新:

发表机构

Robotics Institute, Carnegie Mellon University(卡内基梅隆大学机器人研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对长视频理解受限于上下文窗口的问题,提出无需训练的关键帧采样框架LENS,它能基于文本查询动态分配帧预算,在时空缩放间平衡,提升关键帧采样效果,优于现有方法,提高了Video-MME准确率。

AI 中文摘要

尽管多模态大语言模型取得了快速进展,但长视频理解仍受限于有限的上下文窗口。近期关键帧采样方法试图通过将视频输入提炼为紧凑的查询相关帧集来缓解此问题,但在广阔的时空搜索空间中导航仍具挑战,因为空间细节和时间覆盖常相互冲突。为解决此问题,我们引入LENS,这是一个无需训练的关键帧采样框架,它基于文本查询动态决定何时放大获取细粒度细节以及何时缩小获取更广泛上下文。具体而言,LENS在空间放大(突出单个帧内查询相关区域)和时间缩小(通过多帧聚合扩展时间范围)之间自适应分配有限的帧预算,使模型能够在多个粒度上进行推理,同时捕获高保真细节和长距离上下文。在各种长视频基准测试中,LENS始终优于先前的先进关键帧采样方法,并比均匀采样有显著提升,将Video-MME准确率从53.3%提高到60.7%。

英文摘要

Despite rapid progress in Multi-modal Large Language Models (MLLMs), understanding long-form videos is still bottlenecked by limited context windows. While recent keyframe sampling methods attempt to mitigate this by distilling video inputs into a compact set of query-relevant frames, navigating the vast spatio-temporal search space remains challenging, as spatial detail and temporal coverage often conflict. To address this, we introduce LENS, a training-free keyframe sampling framework that dynamically decides when to zoom in for fine-grained details and when to zoom out for broader context based on the text query. Concretely, LENS adaptively allocates a limited frame budget between spatial zoom-ins, which highlight query-relevant regions within individual frames, and temporal zoom-outs, which expand the temporal scope through multi-frame aggregation, enabling the model to reason across multiple granularities while capturing both high-fidelity details and long-range context. Across diverse long-form video benchmarks, LENS consistently outperforms prior state-of-the-art keyframe sampling methods and delivers substantial gains over uniform sampling, improving Video-MME accuracy from 53.3% to 60.7% with Qwen2.5-VL.Code is available at https://github.com/zhangce01/LENS.

CommentsAccepted at ECCV 2026. Project page: https://zhangce01.github.io/LENS/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑