VI3:利用惯性线索实现预训练3D基础模型的度量级锚定
VI3: Grounding Pretrained 3D Foundation Models with Inertial Cues
浏览论文内容
中文总结 AI 辅助
VI3是一种仅用IMU读数即可对预训练3DFM进行度量级锚定的模型无关框架,在合成与真实航空数据集上无需真值监督即可恢复度量尺度,兼具几何一致性,可作为精细优化或强先验。
中文摘要 AI 辅助
3D基础模型(3DFMs)擅长从场景的多视角图像预测相机位姿和稠密深度,展现出强大的零样本泛化能力。但由于单目图像无法观测度量尺度,其绝对尺度预测通常不准确。大多数设备配备的惯性测量单元(IMUs)可通过观测带尺度的运动,自然补充单目相机的不足。我们提出VI3,一种模型无关的框架,仅利用IMU读数即可对预训练3DFM进行度量级锚定。VI3初始化并预积分IMU以获取度量运动参考,再用该参考恢复3DFM输出的尺度。我们的方法包含适配不同3DFM架构的可调整锚定策略。在合成数据集和真实航空数据集上的实验表明,VI3无需真值监督即可恢复度量尺度,同时保持几何一致性;在运动条件良好时作为精细优化,在运动信息不足时作为强先验。
英文摘要
3D foundation models (3DFMs) excel at predicting camera poses and dense depth from multiple views of a scene, showcasing strong zero-shot generalization. However, as metric scale is not observable from monocular images, their absolute scale predictions are typically inaccurate. Inertial measurement units (IMUs), present in most devices, naturally complement monocular cameras by observing scaled motion. We introduce VI3, a model-agnostic framework that metrically anchors a pretrained 3DFM using only IMU readings. VI3 initializes and preintegrates the IMU to obtain a metric motion reference, which is then used to recover the scale of the 3DFM outputs. Our method includes adaptable anchoring strategies tailored to diverse 3DFM architectures. Experiments on synthetic and real aerial datasets demonstrate that VI3 recovers metric scale without ground-truth supervision while preserving geometric consistency, acting as a fine refinement under well-conditioned motion and as a strong prior when motion is less informative.
发表机构
- Universidad de Zaragoza(萨拉戈萨大学)
机构由 AI 辅助整理,请以论文原文为准。