发表机构
Kalinga Institute of Industrial Technology (KIIT); University of Southern California(卡林加工业技术学院; 南加州大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
LazySloth提出有界惰性树搜索,利用VLM对不相关视频部分进行有界描述,实现2.9-8.3倍加速,并在多个基准上保持或提升准确性,缩小开源与闭源模型差距。
AI 中文摘要
现代视觉语言模型(VLM)在长视频理解方面已展现出有前景的结果,这得益于它们能够捕获丰富的语义信息。然而,大多数方法侧重于对提取的图像帧进行粗粒度描述,这在计算上效率低下,并且需要具有大上下文窗口的模型。尽管已有工作通过多模态检索增强生成(RAG)探索了高效方法,但它们依赖于有损嵌入,从而丢失了时间上下文和细粒度细节。迄今为止,很少有工作研究如何优化基于VLM的查询相关信息检索。我们引入了LazySloth,一种高效的基于树的搜索方法,通过对视频中被VLM视为不相关的部分进行有界描述,将视频理解和检索任务加速2.9-8.3倍(与现有智能体方法相比)。与当代专门的视频理解VLM和基于RAG的方法相比,LazySloth在两个最近的开源基础VLM——Gemma 4 31B和Qwen3.6 27B——上,在四个基准测试中取得了相似或更好的最终任务准确性。LazySloth缩小了基础开源模型与闭源模型GPT-4o之间的差距。消融实验表明,用基于CLIP的检索替换VLM场景理解会导致8.8-19.9%的准确性损失,而惰性树构建在描述成本的一小部分下即可匹配急切构建。通过LazySloth,我们证明了在不显著损失性能的情况下实现更快长视频理解的可能性。
英文摘要
Modern vision-language models (VLMs) have shown promising results in long-video understanding due to the rich semantic information they can capture. However, most methods focus on coarse captioning of extracted image frames that are computationally inefficient and require models with large context windows. While past work has explored efficient methods through multimodal retrieval-augmented generation (RAG), they rely on lossy embeddings that lose temporal context and fine-grained detail. Few works to date have investigated how VLM-based query-relevant information retrieval can be optimized. We introduce LazySloth, an efficient tree-based search method that speeds up video comprehension and retrieval tasks 2.9-8.3x (compared to existing agentic methods) through bounded captioning of portions of the video considered irrelevant by a VLM of the video. Compared to contemporary specialized video-understanding VLMs and RAG-based methods, LazySloth achieved similar or better final task accuracy across two recent open-source base VLMs--Gemma 4 31B and Qwen3.6 27B--across four benchmarks. LazySloth reduced the gap between the base open-source model and a closed-source model, GPT-4o. Ablations showed that replacing VLM scene understanding with CLIP-based retrieval cost 8.8-19.9% in accuracy, while lazy tree construction matches eager construction at a fraction of the captioning cost. With LazySloth, we demonstrate the possibility of faster long-video comprehension without substantial loss in performance.
CommentsUnder review at conference. Preprints allowed when under review