StreamTTT:在流式视觉语言模型中协调实时感知与长期记忆
StreamTTT: Reconciling Real-Time Perception and Long-Term Memory in Streaming VLMs
浏览论文内容
中文总结 AI 辅助
StreamTTT通过将长程历史写入注意力外的快速权重、保留短滑动键值缓存,在OVO-Bench和StreamingBench上提升了流式VLMs的实时感知与长程记忆能力,且代码将公开。
中文摘要 AI 辅助
人类能轻松感知当下并回忆过去,但流式视觉语言模型(streaming VLMs)常需在实时感知与长期记忆间做权衡。现有研究表明,缩短上下文可提升当前场景感知能力,但会牺牲长程回忆。为协调这两种能力,我们提出StreamTTT,它将长程历史写入注意力上下文之外的在线更新快速权重中,仅保留短的滑动键值缓存用于近期证据,缓解注意力稀释问题。我们在离线长视频问答和新构建的实时问答语料库上联合训练StreamTTT。在OVO-Bench上,StreamTTT-4B的实时感知比SimpleStream-4B高1.4个点,反向追踪比其高3.7个点;在StreamingBench的实时视觉理解(RTVU)子集上,它与更大规模的SimpleStream-8B表现相当。我们将发布代码。
英文摘要
Humans effortlessly perceive the present while remembering the past, yet streaming VLMs often trade off real-time perception against long-term memory. Prior work shows that shortening the context can sharpen current-scene perception at the expense of long-range recall. To reconcile these abilities, we introduce StreamTTT, which writes long-range history into online-updated fast weights outside the attention context. This leaves a short sliding key-value cache dedicated to recent evidence, mitigating attention dilution. We train StreamTTT jointly on offline long-video QA and a newly constructed real-time QA corpus. On OVO-Bench, under each model's reported input protocol, StreamTTT-4B outperforms the same-scale SimpleStream-4B by 0.6 points in real-time perception and 5.3 points in backward tracing. It also surpasses the larger SimpleStream-8B by 0.73 points on StreamingBench's Real-Time Visual Understanding (RTVU) subset. Our code is publicly available at https://github.com/zeyun-zhong/StreamTTT.
发表机构
- National University of Singapore(新加坡国立大学)
- Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)
机构由 AI 辅助整理,请以论文原文为准。