发表机构
The University of Hong Kong; Voyager Research, DiDi Chuxing(香港大学; 滴滴出行探索研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FoundationGeo是两阶段框架,第一阶段学习高保真几何模型,第二阶段引入校准场生成3D点图。它还解决相机固有覆盖问题,通过合成数据提高鲁棒性。在多基准零射击评估中,该方法增强跨域鲁棒性,性能优于其他基线。
AI 中文摘要
我们提出了FoundationGeo,这是一个两阶段框架,通过空间校准和有原则的数据设计明确地连接相对和度量预测。第一阶段通过用DINOv3初始化并在精心策划的1020万个样本的多域语料库上进行训练,学习一个高保真、仿射不变的几何模型,并辅以局部细节监督,产生清晰的边界和强大的数据泛化能力。第二阶段通过引入用于度量估计的轻量级逐像素校准场超越全局缩放:一个用于空间变化度量对齐的比例场和一个减轻点图几何中方向偏差的光线方向校正场,共同生成度量一致的3D点图。除了模型设计,我们还将相机固有覆盖范围,特别是训练和测试数据之间的焦距分布不匹配,识别为零射击度量泛化的关键瓶颈:当测试固有值落在训练分布之外时,性能会急剧下降。为了解决这个问题,我们使用基于Blender的数据引擎合成跨不同焦距的额外训练数据,修复覆盖不足的焦距区域并提高固有偏移下的鲁棒性。在七个基准上进行的广泛零射击评估表明,FoundationGeo显著增强了跨域鲁棒性,在不同领域中保持在前列,同时避免了其他方法中观察到的急剧跨域性能下降。这种一致性转化为最佳的整体性能,平均比更重的基线高出5.2%以上。
英文摘要
We present FoundationGeo, a two-stage framework that explicitly bridges relative and metric prediction via spatial calibration and principled data design. Stage 1 learns a high-fidelity, affine-invariant geometry model by initializing with DINOv3 and training on a curated 10.2M-sample multi-domain corpus with complementary local-detail supervision, yielding sharp boundaries and strong cross-domain generalization. Stage 2 moves beyond global scaling by introducing lightweight pixel-wise calibration fields for metric estimation: a scale field for spatially varying metric alignment and a ray-direction correction field that mitigates directional bias in point-map geometry, together producing metrically consistent 3D point maps. Beyond model design, we identify camera intrinsic coverage, especially focal length distribution mismatch between training and test data, as a key bottleneck for zero-shot metric generalization: performance drops sharply when test intrinsics fall outside the training distribution. To address this, we synthesize additional training data across diverse focal lengths using a Blender-based data engine, repairing under-covered focal regimes and improving robustness under intrinsic shift. Extensive zero-shot evaluations across seven benchmarks show that FoundationGeo significantly strengthens cross-domain robustness, staying near the top across diverse domains while avoiding the sharp cross-domain performance drops observed in other methods. This consistency translates into the best overall performance, surpassing heavier baselines by over 5.2% on average.
CommentsAccepted to ECCV 2026. Project page: https://mx-liu6.github.io/FoundationGeo-web/