发表机构
University of California, San Diego; The University of Sydney; Central South University; Harbin Institute of Technology (Shenzhen); SenseTime Research(加利福尼亚大学圣地亚哥分校; 悉尼大学; 中南大学; 哈尔滨工业大学(深圳); 商汤科技研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
StateTrace是一种以对象为中心的框架,为VideoLLMs赋予长视频隐状态推理能力,构建时空状态记忆并经实验显著提升了VideoLLMs在相关基准上的性能。
AI 中文摘要
现有的视觉语言模型(VLMs)在视频理解方面已取得出色性能,但当目标对象不可见时,它们在长视频时空推理任务中仍存在困难,常将“不可见”误判为“未知”。我们将这一挑战定义为隐状态时空推理:即从上下文交互中推断对象在长时间不可见区间内的状态。为解决该问题,我们提出StateTrace,这是一种新型以对象为中心的框架,为视频语言模型(VideoLLMs)赋予长视频中隐状态推理的显式机制。StateTrace构建了可复用的时空状态记忆,将对象轨迹、对象间关系及状态转换事件组织为结构化推理基底。推理时,它会检索与问题相关的状态演化轨迹并将其转换为紧凑推理线索,使模型能显式推断对象消失的原因、其在不可见期间的状态如何演化,以及该状态在查询时刻是否应保留。我们还构建了HSR-Bench,这是一个用于隐状态推理的诊断基准,包含来自1384个独特视频的1427个视频问答样本。对多个VideoLLMs的大量实验表明,StateTrace在公开基准和HSR-Bench上均能持续提升性能,例如将VideoLLaMA3在HSR-Bench上的表现从39.6提升至64.2。
英文摘要
Existing VLMs have achieved strong performance in video understanding, yet they struggle with long-video spatiotemporal reasoning when target objects become invisible, often mistaking "invisible" for "unknown". We define this challenge as hidden-state spatiotemporal reasoning: inferring object states during prolonged invisible intervals from context interactions. To address this, we propose StateTrace, a novel object-centric framework that endows VideoLLMs with an explicit mechanism for hidden state reasoning in long videos. StateTrace builds a reusable spatiotemporal state memory that organizes object trajectories, inter-object relations, and state-transition events into a structured reasoning substrate. At inference time, it retrieves question-relevant state-evolution trajectories and converts them into compact reasoning cues, enabling the model to explicitly reason about why an object disappears, how its state evolves while invisible, and whether that state should persist at query time. We further build HSR-Bench, a diagnostic benchmark for hidden-state reasoning, containing 1,427 video-QA samples from 1,384 unique videos. Extensive experiments across multiple VideoLLMs show that StateTrace consistently improves performance on both public benchmarks and HSR-Bench (e.g., improving VideoLLaMA3 from 39.6 to 64.2 on HSR-Bench).
Comments10 pages. Accepted at ACM Multimedia 2026 (ACM MM 2026)