发表机构
Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多模态大语言模型处理流视频计算成本高的问题,提出ShallowStream框架,通过利用MLLM浅层构建索引并检索证据,在保持性能的同时大幅降低了延迟。
AI 中文摘要
流视频理解是具身智能、自动驾驶、工业监控、预警系统、可穿戴助手等现实应用的关键能力。然而,用多模态大语言模型(MLLM)处理连续视频流的计算成本极高。现有研究已通过视觉令牌剪枝、令牌合并、量化、按需帧检索、上下文卸载等方式降低流处理开销,但多数方法忽略了模型深度维度:对传入帧重复执行全深度MLLM预填充的成本过高,会产生大量计算开销,且KV缓存会随预填充深度成正比增长。为解决这些挑战,我们提出ShallowStream这一新框架,利用MLLM的浅层同时执行帧编码和检索索引构建。流处理期间,ShallowStream通过浅层的KV缓存维持一个始终在线的轻量索引;查询阶段回答时,我们利用浅层生成的注意力分数对上下文帧打分,并采用感知多样性的选择策略检索精确且全面的证据。ShallowStream的性能与现有最强流方法相当,同时将每帧预填充延迟和10秒端到端延迟分别降低了最高52.1倍和11.9倍。我们的代码可在此https URL获取。
英文摘要
Streaming video understanding is a critical capability for real-world applications, including embodied intelligence, autonomous driving, industrial monitoring, surveillance and early warning, and wearable assistants. However, processing continuous video streams with multimodal large language models (MLLMs) is computationally expensive. Existing efforts have explored reducing streaming overhead through visual token pruning, token merging, quantization, on-demand frame retrieval, and context offloading. However, most existing methods overlook the dimension of model depth. Repeatedly executing full-depth MLLM prefill over incoming frames is prohibitively expensive, incurring substantial computational overhead and causing the KV cache to grow at a rate directly proportional to the prefill depth. To address these challenges, we propose ShallowStream, a novel framework that leverages the shallow layers of an MLLM to simultaneously perform frame encoding and retrieval index building. During stream processing, ShallowStream maintains an always-on lightweight index using the KV cache of shallow layers. During query-time answering, we leverage the attention scores generated by the shallow layers to score context frames and employ a diversity-aware selection strategy to retrieve precise and comprehensive evidence. ShallowStream achieves performance on par with the strongest existing streaming methods, while reducing per-frame prefill latency and 10-second end-to-end latency by up to 52.1x and 11.9x, respectively. Our code is available at https://github.com/CURRENTF/ShallowStream.