TempoGround:基于视觉语言模型的状态感知流可视化定位
TempoGround: State-Aware Streaming Visual Grounding with Vision-Language Models
浏览论文内容
中文总结 AI 辅助
TempoGround是基于VLM的框架,通过状态感知跨帧对应与SGR强化解决流可视化定位的身份漂移等问题,在多基准上显著提升了2D、3D定位指标。
中文摘要 AI 辅助
可视化定位将语言指代映射到空间目标,是基于视觉语言模型(VLM)的开放词汇感知的核心。现有方法在单帧和视频可视化定位上已取得显著进展,但在流输入下仍存在身份漂移、跨帧不一致性以及部分遮挡下定位脆弱的问题。为解决这些问题,我们提出了TempoGround,一种原生基于VLM的框架,用于检测跨帧物体对应关系并显式建模物体存在状态,从而在流输入下实现准确、一致的可视化定位。其核心是由状态感知跨帧对应关系引导的课程预测机制:TempoGround解决二维实例关联,预测每个物体是新进入、持续存在还是离开视野,解码二维边界框,再将其提升为相机坐标系下的三维边界框。由于仅令牌级监督无法捕捉流定位的几何目标,我们进一步引入流定位强化(Streaming Grounding Reinforcement, SGR),通过可验证的定位、身份和一致性奖励优化TempoGround,共同强化持久定位和时间一致的预测。我们精心设计了三阶段训练策略,并在大规模数据上训练TempoGround。我们在多个具有挑战性的基准上评估了因果流输入下的可视化定位:TempoGround平均将F1_2D@0.5和F1_2D@0.95分别提升4.4和0.5,F1_3D@0.25和AP_3D分别提升6.2和7.5。这些结果表明,TempoGround为流输入下的可视化定位提供了实用基础。
英文摘要
Visual grounding maps language referents to spatial targets and is central to open-vocabulary perception with vision-language models. Existing methods have made substantial progress on single-frame and video-based visual grounding, yet under streaming inputs they still suffer from identity drift, cross-frame inconsistency, and fragile localization under partial occlusion. To address these issues, we present TempoGround, a VLM-native framework that detects cross-frame object correspondence and explicitly models object presence states, thereby enabling accurate and consistent visual grounding under streaming inputs. The key is a curriculum prediction mechanism guided by state-aware cross-frame correspondence: TempoGround resolves 2D instance association, predicts whether each object newly enters, continues in, or leaves the view, decodes the 2D box, and then lifts it to a camera-frame 3D box. As token-level supervision alone cannot capture the geometric objectives of streaming grounding, we further introduce Streaming Grounding Reinforcement (SGR), which optimizes TempoGround with verifiable Grounding, Identity, and Consistency rewards, jointly reinforcing persistent localization and temporally consistent predictions. We carefully design a three-stage training strategy and train TempoGround on large-scale data. We evaluate visual grounding under causally streaming inputs on multiple challenging benchmarks: TempoGround improves F1_2D@0.5 and F1_2D@0.95 by 4.4 and 0.5 on average, and F1_3D@0.25 and AP_3D by 6.2 and 7.5, respectively. These results demonstrate that TempoGround provides a practical foundation for visual grounding under streaming inputs.
发表机构
- EngineAI
机构由 AI 辅助整理,请以论文原文为准。