arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ChronoStitch:用于长时程时间推理的视觉键值记忆的免训练组合

ChronoStitch: Training-Free Composition of Visual KV Memories for Long-Horizon Temporal Reasoning

Santiram Tiwari, Nishant Sinha, Kunal Kislay

arXiv 2607.19547首次发表:更新:

发表机构

KGraph AI Solutions Pvt. Ltd.(KGraph人工智能解决方案私人有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对长视频问答中视觉KV记忆组合问题,ChronoStitch提出免训练方法,先将键值重基于全局坐标系,解决几何不一致问题,再处理残余内容差距,通过选择性重新计算部分视觉令牌,提升事件排序准确性且运行速度更快。

AI 中文摘要

长视频问答要求模型长时间保留视觉证据,而无需重复重新处理同一视频。一种实用方法是为每个视频块存储视觉语言模型的内部键值(KV)缓存,并在查询时检索该状态。然而,独立缓存的视频块无法正确组合:每个块都从本地旋转位置零预填充,因此简单的连接会冲突时间相位并消除有关首先发生了什么、事件发生频率或视频中发生了什么变化的问题所需的全局顺序。本文提出了ChronoStitch,一种用于组合独立存储的视觉KV记忆的免训练方法。该方法首先将存储的旋转后键重新基于全局三轴多模态RoPE坐标系,该坐标系保留时间、高度和宽度结构。我们展示了为什么一维标量重新索引对于视觉令牌在几何上是不一致的,因为它将帧内的空间顺序转换为错误的时间位移。然后,我们解决了位置修复留下的残余内容差距:后来的块最初是在不关注早期块的情况下编码的。因此,ChronoStitch选择性地重新计算一小部分高偏差的后期块视觉令牌,同时允许它们关注组合缓存。在Qwen2.5-VL-3B和TempCompass的时间分割上,ChronoStitch优于简单组合和仅位置变体,提高了事件排序准确性,同时运行速度比完全联合重新预填充快3.3倍。

英文摘要

Long-video question answering requires a model to preserve visual evidence over time without repeatedly reprocessing the same video. A practical approach is to store the vision-language model's internal key-value (KV) cache for each video chunk and retrieve that state at query time. However, independently cached video chunks do not compose correctly: every chunk is prefilled from local rotary position zero, so naive concatenation collides temporal phases and removes the global order required for questions about what happened first, how often events occurred, or what changed across the video. This paper presents ChronoStitch, a training-free method for composing independently stored visual KV memories. The method first re-bases stored post-rotary keys onto a global three-axis multimodal RoPE coordinate system that preserves time, height, and width structure. We show why a one-dimensional scalar re-indexing is geometrically inconsistent for visual tokens because it turns spatial order within a frame into false temporal displacement. We then address the residual content gap left by positional repair: later chunks were originally encoded without attending to earlier chunks. ChronoStitch therefore selectively recomputes a small fraction of high-deviation later-chunk visual tokens while allowing them to attend over the composed cache. On Qwen2.5-VL-3B and the temporal split of TempCompass, ChronoStitch outperforms naive composition and position-only variants, improving event-ordering accuracy while running 3.3x faster than full joint re-prefilling.

Comments6 pages, 4 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑