arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29881cs.CV

OptiGeo:面向光学挑战性场景具身感知的高效单目几何方法

OptiGeo: Efficient Monocular Geometry for Embodied Perception in Optically Challenging Scenes

Muxin Liu, Tianbo Liu, Jing Xia, Xiaoyang Lyu, Xiaoshan Wu, Bo Wang, Peng Dai, Zhongrui Wang, Shaoshuai Shi, Xiaojuan Qi

首次发表
浏览论文内容

中文总结 AI 辅助

OptiGeo是一种感知偏差的训练框架,利用小目标渲染集修正局部几何畸变,在透明场景基准上性能远超更大规模模型,可作为光学挑战性场景的高效感知模块。

中文摘要 AI 辅助

单目深度估计已实现强大的开放域泛化,但在透明、反射和镜面环境中,可靠的机器人部署仍存在困难,这类环境中深度传感器常产生缺失或有偏差的深度数据。现有方法通常通过特定场景的预处理、辅助模块或事后微调来处理这类光学故障,尽管在受限场景中有效,但这些设计增加了架构冗余,且可能过度将通用几何模型特化为狭窄的光学场景。我们将该问题重新定义为基础模型训练中的局部故障模式,并确定传感器诱导的监督偏差是关键瓶颈:模型从光学挑战性区域的有偏差真实深度监督中继承了传感器故障模式。随后我们引入OptiGeo,这是一种感知偏差的训练框架,它利用干净几何教师和残差修剪对齐来修复有偏差的真实监督。我们将针对透明目标的渲染重新定义为紧凑的干净光学几何源,而非大型特定领域微调集。仅用小的目标渲染集,OptiGeo就能学习透明物体和区域的几何结构,修正真实传感器无法可靠监督的局部几何畸变。尽管仅有3000万参数,OptiGeo在透明场景基准上的性能远超规模大得多的3亿级单目模型和数十亿级多视图基线,同时在通用零样本深度和边界清晰度上仍具竞争力。实际导航案例进一步验证了它作为光学挑战性场景中高效感知模块的实用性。

英文摘要

Monocular depth estimation has achieved strong open-domain generalization, yet reliable robotic deployment remains difficult in transparent, reflective, and specular environments, where depth sensors often produce missing or biased depth. Existing methods often handle such optical failures with scene-specific preprocessing, auxiliary modules, or post-hoc fine-tuning. While effective in constrained settings, these designs increase architectural redundancy and can over-specialize general geometry models to narrow optical scenarios. We revisit this problem as a localized failure mode within base-model training and identify sensor-induced supervision bias as a key bottleneck: models inherit sensor failure patterns from biased real-depth supervision in optically challenging regions. We then introduce OptiGeo, a bias-aware training framework that rehabilitates biased real supervision using a clean-geometry teacher and residual-trimmed alignment. We redefine transparency-targeted rendering as a compact source of clean optical geometry, rather than a large domain-specific fine-tuning set. With only a small targeted rendering set, OptiGeo learns the geometric structure of transparent objects and regions, correcting local geometry distortions that real sensors cannot reliably supervise. Despite only 30M parameters, OptiGeo outperforms substantially larger 300M-scale monocular models and billion-scale multi-view baselines on transparent-scene benchmarks, while remaining competitive on general zero-shot depth and boundary sharpness. Real-world navigation cases further validate its practicality as an efficient perception module in optically challenging scenes.

发表机构

  • The University of Hong Kong(香港大学)
  • Voyager Research, DiDi Chuxing(滴滴出行 Voyager 研究院)
  • Southern University of Science and Technology(南方科技大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑