TRIG:轨迹-装置解耦度量几何学习
TRIG: Trajectory-Rig Decoupled Metric Geometry Learning
查看机构详情
- Carizon(卡里松)
- ShanghaiTech University(上海科技大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对自动驾驶中视觉几何模型不适用于刚性多相机驱动系统的问题,提出TRIG框架,将相机姿态分解,引入解耦编码与监督及稀疏时空注意力,在多基准测试中实现姿态估计、深度预测和3D重建的最优性能。
中文摘要 AI 辅助
以视觉为中心的自动驾驶需要从同步多相机观测中进行精确的度量几何和自我运动估计。近期视觉几何模型在姿态估计、深度预测和3D重建方面表现出色,但不适用于刚性多相机驱动系统。它们常将相机姿态编码为纠缠表示,联合建模时变自我运动和静态相机装置几何,限制了车辆侧几何先验的利用。我们提出轨迹-装置解耦度量几何学习(TRIG),将相机姿态分解为自我轨迹和相机装置组件,实现自我运动和静态多相机拓扑的单独建模。引入解耦姿态编码和监督,为度量一致学习分别约束轨迹演化和装置几何。此外,稀疏时空注意力将跨相机交互与时间聚合分离,降低全局注意力成本同时保留几何推理。在五个自动驾驶基准测试上的实验表明,TRIG在姿态估计、度量深度预测和3D重建方面达到了当前最优性能。
英文摘要
Vision-centric autonomous driving requires accurate metric geometry and ego-motion estimation from synchronized multi-camera observations. Recent visual geometry models show strong performance in pose estimation, depth prediction, and 3D reconstruction, but are not tailored to rigid multi-camera driving systems. They often encode camera poses as entangled representations, in which time-varying ego-motion and static camera-rig geometry are jointly modeled, limiting the utilization of vehicle-side geometric priors. We propose Trajectory-Rig Decoupled Metric Geometry Learning (TRIG), a geometry perception framework for autonomous driving. TRIG factorizes camera poses into ego-trajectory and camera-rig components, enabling separate modeling of ego-motion and static multi-camera topology. We introduce decoupled pose encoding and supervision, which separately constrain trajectory evolution and rig geometry for metric-consistent learning. Moreover, sparse Temporal--Spatial attention separates cross-camera interaction from temporal aggregation, reducing global attention cost while preserving geometric reasoning. Experiments on five autonomous driving benchmarks show that TRIG achieves state-of-the-art performance in pose estimation, metric depth prediction, and 3D reconstruction.