arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.19228cs.CV

IGGT4D:流式4D实例关联几何变换器

IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer

Zhengyu Zou, Hao Li, Kuixuan Jiao, Liu Liu, Tingyang Xiao, Xiaolin Zhou, Fangzhou Hong, Zhizhong Su, Dingwen Zhang, Ziwei Liu

首次发表
浏览论文内容

中文总结 AI 辅助

针对现实世界空间智能中视频流场景理解问题,提出IGGT4D流式实例关联几何变换器,通过因果时空建模更新统一表示,还构建数据集。实验证明其在长动态序列在线推理上优于现有基线。

中文摘要 AI 辅助

现实世界的空间智能要求智能体从连续视频流中理解场景,物体随时间移动、持续、消失和再现。近期空间基础模型实现了可泛化的前馈3D重建,但多数流式方法仍以几何为中心,缺乏时间一致的对象级理解。现有语义重建和3D感知视觉语言方法依赖外部提取的2D语义线索或松散耦合的几何输入。本文提出IGGT4D,一种用于在线4D场景理解的流式实例关联几何变换器,通过因果时空建模重用历史上下文,增量更新相机运动、几何和对象身份的统一表示。为解决高质量4D监督缺失问题,构建了InsScene4D - 147K数据集。实验表明IGGT4D优于现有流式基线,能对长动态序列进行可扩展的在线推理。

英文摘要

Real-world spatial intelligence requires agents to understand scenes from continuous video streams, where objects move, persist, disappear, and reappear over time. While recent spatial foundation models have enabled generalizable feed-forward 3D reconstruction, most streaming methods remain geometry-centric and lack temporally consistent object-level understanding. Meanwhile, existing semantic reconstruction and 3D-aware vision-language methods largely rely on externally extracted 2D semantic cues or loosely coupled geometry inputs, limiting unified geometry-instance learning in long dynamic scenes. In this paper, we propose IGGT4D, a streaming instance-grounded geometry Transformer for online 4D scene understanding. IGGT4D processes video frames sequentially, reuses historical context through causal spatial-temporal modeling, and incrementally updates a unified representation of camera motion, geometry, and object identity. This enables long-sequence feed-forward reconstruction with geometry-instance consistency in dynamic environments. To address the lack of high-quality 4D supervision, we further construct InsScene4D-147K, a large-scale dataset spanning real/synthetic and static/dynamic scenes, with RGB images, depth, poses, and temporally consistent instance masks generated by an automated geometry-guided annotation pipeline. Experiments on 3D reconstruction, pose estimation, instance spatial tracking, and open-vocabulary segmentation demonstrate that IGGT4D outperforms existing streaming baselines while maintaining scalable online inference for long dynamic sequences.

发表机构

  • Horizon Robotics(地平线机器人公司)
  • S-Lab, Nanyang Technological University(南洋理工大学S-Lab)
  • Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑