发表机构
The Hong Kong University of Science and Technology (Guangzhou); Beta Infinity; Tsinghua University(香港科技大学(广州); Beta Infinity; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PREM提出一种无需记忆令牌的循环记忆框架,通过分离视频摄取与查询回答,在冻结VLM上实现高效长视频理解,显著提升多基准准确率。
AI 中文摘要
长视频理解必须在严格的令牌预算下捕获瞬态视觉证据,然而现有方法会压缩帧、附加记忆令牌或改变内部键值(KV)缓存。我们引入了前缀引导的循环记忆(PREM),这是一个无需记忆令牌的框架,适用于冻结的视觉语言模型(VLM)。PREM将视频摄取与查询回答分离:一个循环写入器将视觉流蒸馏为紧凑的256 KiB多槽关联状态,而一个基于问题的读出器在预填充期间向现有的非视觉提示前缀添加记忆派生的键/值(K/V)引导调制。这实现了写一次、查询多次的推理,无需额外的提示令牌或解码循环。在六个长视频基准测试中,无论是在离线还是流式端到端设置下,PREM在每个评估的视觉预算下都始终优于冻结基线。在16帧的受限预算下,PREM在Qwen2.5-VL-3B上将宏平均准确率提高了3.06%,其中动作反义词识别提高了11.0%,局部针检索提高了9.9%。这些增益仅需调整0.24%的骨干参数,峰值GPU内存开销为0.03 GiB。
英文摘要
Long-video understanding must capture transient visual evidence under strict token budgets, yet existing methods compress frames, append memory tokens, or alter internal key-value (KV) caches. We introduce Prefix-Steered Recurrent Memory (PREM), a memory-token-free framework for frozen vision-language models (VLMs). PREM separates video ingestion from query answering: a recurrent writer distills visual streams into a compact 256 KiB multi-slot associative state, while a question-conditioned readout adds memory-derived key/value (K/V) steering modulations to existing non-visual prompt prefixes during prefill. This enables write-once, query-many inference without extra prompt tokens or decoding recurrence. Across six long-video benchmarks in offline and streaming end-of-stream settings, PREM consistently outperforms frozen baselines at every evaluated visual budget. Under a constrained budget of 16 frames, PREM improves macro-average accuracy by 3.06% on Qwen2.5-VL-3B, with gains of 11.0% on action antonym identification and 9.9% on localized needle retrieval. These gains require tuning 0.24% of backbone parameters at 0.03 GiB of peak GPU memory overhead.
Comments22 pages, 11 figures, 12 tables