AI 中文总结
StreamRig提出冻结并流式框架,利用多相机刚性几何在冻结3D基础模型上构建因果流式里程计,通过两阶段训练实现零样本真实世界评估,在四个数据集上降低漂移并保持低推理成本。
AI 中文摘要
移动机器人和车辆搭载同步多相机刚性结构,然而许多流式3D基础模型专为单目输入设计,高效利用刚性结构几何仍是一大挑战。我们提出StreamRig,一种冻结并流式框架,在冻结的多视角3D基础模型上为标定刚性结构构建因果流式里程计。冻结前端利用刚性标定联合感知同步视图。刚性重采样器压缩其特征,因果桥接器应用带键值缓存的因果注意力,轻量级头部回归刚性姿态。周期性重新锚定协议支持长序列中的稳定姿态估计。仅训练这些模块,总计7460万参数,以相对姿态作为唯一监督。我们的两阶段训练策略结合组重定位预训练与因果刚性训练,将冻结前端的几何先验和预训练模块的对齐能力迁移至流式里程计。我们在NCLT、TartanGround、KITTI-360以及自采集的人形机器人数据集ZJH上评估,训练仅使用仿真,真实世界评估为零样本。在所有四个数据集上,StreamRig的平移和旋转漂移均低于评估的非神谕单目流式和刚性感知离线模型,同时保持低推理成本。消融实验和控制相机数量实验识别了这些增益的来源。我们进一步研究更长训练窗口对更长时域推理的影响。代码已发布于该https URL。
英文摘要
Mobile robots and vehicles carry synchronized multi-camera rigs, yet many streaming 3D foundation models are designed for monocular input, leaving efficient use of rig geometry a challenge. We present StreamRig, a freeze-and-stream framework that builds causal streaming odometry for calibrated rigs on a frozen multi-view 3D foundation model. The frozen front-end jointly perceives the synchronized views using rig calibration. A Rig-Resampler compresses their features, a CausalBridge applies causal attention with a key-value cache, and a lightweight head regresses rig poses. A periodic re-anchoring protocol supports stable pose estimation over long sequences. Only these modules are trained, 74.6M parameters in total, with relative poses as the sole supervision. Our two-stage training strategy combines group relocalization pretraining with causal rig training to transfer the geometric priors of the frozen front-end and the alignment ability of the pretrained modules to streaming odometry. We evaluate on NCLT, TartanGround, KITTI-360, and our self-collected humanoid-robot dataset ZJH, where training uses only simulation and real-world evaluation is zero-shot. Across all four datasets, StreamRig achieves lower translation and rotation drift than the evaluated non-oracle monocular streaming and rig-aware offline models, while maintaining low inference cost. Ablations and controlled camera-count experiments identify the sources of these gains. We further examine how longer training windows affect inference over longer horizons. Code has been released at https://github.com/WeiYuFei0217/StreamRig.
Comments8 pages, 4 figures, 5 tables. Code: https://github.com/WeiYuFei0217/StreamRig