SpaTime:面向时空推理的流式视觉语言模型
SpaTime: Streaming Vision-Language Models for Spatio-temporal Reasoning
浏览论文内容
中文总结 AI 辅助
SpaTime是一种流式视觉语言模型,通过融合因果几何标记实现时空推理,并采用响应时间损失优化回答时机,在StreamVSTI-Bench上准确率达49.2%,响应时间误差降低66%。
中文摘要 AI 辅助
具身智能体必须在视频仍在到达时对3D空间进行推理,一旦观察到足够的场景就回答问题。融合3D几何先验的视觉语言模型(VLM)实现了强大的空间推理,但它们以离线方式运行,即必须获得完整视频后才能产生答案。流式VLM以因果方式处理帧并自行决定何时响应,但缺乏显式的3D表示。我们提出SpaTime,一种流式VLM,它在每一帧将因果几何标记融合到语言模型中,仅使用迄今为止观察到的帧。为了监督模型何时回答,我们提出了一种响应时间损失,将每帧的响应概率映射为可微的期望响应时间,并惩罚与真实帧的距离。为了评估,我们构建了StreamVSTI-Bench和StreamVSI-Bench,它们是VSTI-Bench和VSI-Bench的流式改编。在StreamVSTI-Bench上,SpaTime达到了49.2%的整体准确率,并将平均响应时间误差相对于最强的流式基线降低了66%。
英文摘要
Embodied agents must reason about 3D space while the video is still arriving, answering questions as soon as they have observed enough of the scene. VLMs that incorporate 3D geometric priors achieve strong spatial reasoning, but they operate offline, i.e., the full video must be available before they produce an answer. Streaming VLMs process frames causally and decide for themselves when to respond, yet they lack explicit 3D representations. We present SpaTime, a streaming VLM that fuses causal geometry tokens into the language model at every frame, using only the frames observed so far. To supervise when the model answers, we propose a response-time loss that maps per-frame response probabilities to a differentiable expected response time and penalizes the distance from the ground-truth frame. For evaluation, we construct StreamVSTI-Bench and StreamVSI-Bench, streaming adaptations of VSTI-Bench and VSI-Bench. On StreamVSTI-Bench, SpaTime reaches 49.2% overall accuracy and reduces the mean response-time error by 66% relative to the strongest streaming baseline.
发表机构
- Purdue University(普渡大学)
- Goertek Alpha Labs(歌尔阿尔法实验室)
机构由 AI 辅助整理,请以论文原文为准。