AI 中文总结
本文提出SlotNarrative,一种用于Video-LLM的紧凑结构化视觉接口,通过持久对象叙事减少视觉令牌数量,在多数据集上实现准确率与令牌数的良好权衡。
AI 中文摘要
视频大语言模型(Video-LLM)在开放式视频理解方面已取得显著进展,但其视觉接口仍存在令牌密集的问题,且缺乏明确结构来关联跨时间的重复对象证据。本文提出SlotNarrative,一种基于槽的接口,将视频组织为由紧凑对象状态令牌表示的持久对象叙事。SlotNarrative并非在建立时间对应关系前压缩逐帧特征,而是先将视觉特征分组为类对象槽,再通过轻量、无参数的内存将重复观测与剪辑级对象条目关联,该内存整合了多种互补匹配线索。每个保留的条目被序列化为两种令牌类型:总结持久对象外观的身份令牌,以及编码片段级外观、几何、可见性和轨迹信息的一组状态令牌。该设计为冻结的Video-LLM仅提供144个分配的视觉令牌位置,与采样帧数无关。在多个数据集上,SlotNarrative相比现有紧凑Video-LLM接口,在准确率与视觉令牌数量间实现了良好权衡,实验结果证实持久对象叙事是一种紧凑、结构化且时间有序的Video-LLM视觉接口,代码将公开提供。
英文摘要
Video large language models (Video-LLMs) have made strong progress in open-ended video understanding. However, their visual interfaces remain token-intensive and provide limited explicit structure for linking recurring object evidence across time. We introduce SlotNarrative, a slot-based interface that organizes a video into persistent object narratives represented by compact object-state tokens. Rather than compressing frame-wise features before establishing temporal correspondence, SlotNarrative first groups visual features into object-like slots and then associates recurring observations with clip-level object entries through a lightweight, parameter-free memory that integrates multiple complementary matching cues. Each retained entry is serialized into two token types: an identity token that summarizes persistent object appearance and a set of state tokens that encode segment-level appearance, geometry, visibility, and trajectory information. This design yields an interface of only 144 allocated visual-token positions for a frozen Video-LLM, independent of the number of sampled frames. Across multiple datasets, SlotNarrative achieves a favorable trade-off between accuracy and visual-token count compared with prior compact Video-LLM interfaces. Experimental results establish persistent object narratives as a compact, structured, and temporally organized visual interface for Video-LLMs. Our code will be made publicly available.