AI 中文总结
针对单目深度基础模型在非朗伯表面易产生深度幻觉的问题,提出无需深度标签的参数高效微调框架 GIFT,通过几何不变性抑制幻觉,在提升镜子与透明物体深度预测的同时保留基础模型性能。
AI 中文摘要
得益于大规模合成训练数据,单目深度基础模型展现出了较强的泛化能力,但它们常于非朗伯表面产生深度幻觉——估计镜子中的反射内容或玻璃后方的透射内容,而非物理表面本身。用真实世界数据适配这些模型颇具挑战,因为传统深度传感器在这类区域也不可靠。我们注意到,非朗伯表面的外观虽随其反射或透射环境变化,但其 underlying 几何保持不变。基于此,我们提出 GIFT(Geometry-Invariant Fine-Tuning,几何不变微调),这是一种无需测量深度标签的参数高效后训练框架。我们在相机与目标几何固定的情况下,采集外观变化受控的 RGB 图像组。GIFT 利用这些观测间的几何不变性,抑制非朗伯深度幻觉,同时保留通用深度估计能力。我们还构建了一个受控基准,用于评估非朗伯深度恢复、外观变化鲁棒性及其他区域的性能保留。在我们的基准及独立真实世界数据集上的实验表明,GIFT 提升了镜子与透明物体的深度预测效果,同时在很大程度上保留了基础模型的性能,为适配单目深度基础模型至非朗伯场景提供了一种实用且低成本的方法。
英文摘要
Monocular depth foundation models, benefiting from large-scale synthetic training data, have demonstrated strong generalization. However, they often hallucinate depth on non-Lambertian surfaces, estimating reflected content in mirrors or transmitted content behind glass rather than the physical surface itself. Adapting these models with real-world data is challenging because conventional depth sensors are also unreliable in such regions. We observe that while the appearance of a non-Lambertian surface varies with its reflected or transmitted environment, its underlying geometry remains unchanged. Based on this observation, we propose GIFT (Geometry-Invariant Fine-Tuning), a parameter-efficient post-training framework that requires no measured depth labels. We collect groups of RGB images under controlled appearance changes while keeping the camera and target geometry fixed. GIFT exploits geometric invariance across these observations to suppress non-Lambertian depth hallucinations while retaining general depth estimation capability. We further construct a controlled benchmark that evaluates non-Lambertian depth recovery, robustness to appearance changes, and performance retention in other regions. Experiments on our benchmark and an independent real-world dataset demonstrate that GIFT improves depth prediction for mirrors and transparent objects while largely preserving the base model's performance, providing a practical and low-cost approach for adapting monocular depth foundation models to non-Lambertian scenes.