AI 中文总结
研究提出X-lens模型用于异构相机实时度量深度估计,基于几何感知异构相机公式,含可学习校准令牌和雅可比参数化失真偏差,在多数据集训练,实验证明其精度高、参数少,在多种设置下性能有竞争力。
AI 中文摘要
我们提出了X-lens,这是一种紧凑的前馈模型,用于从可变数量的校准鱼眼和针孔视图中进行度量深度估计。为了支持实时下游感知,X-lens围绕具有两个关键组件的几何感知异构相机公式构建。可学习的校准令牌在鱼眼和针孔投影空间之间提供粗略对齐,而注入交叉注意力模型的雅可比参数化失真偏差会改变局部投影并促进跨相机一致性,仅用0.04B参数就能实现高达41 FPS的稳健泛化。该模型预测密集深度和全局度量尺度,避免增加计算和优化复杂度的辅助重建目标。为了在尺度和深度上学习这种跨相机泛化,X-lens在多个公共数据集和我们新发布的包含约266K同步六视图帧、170万张单独图像以及103个室内和室外场景的大型合成数据集OmniScene上进行训练。在真实世界和合成室内外数据集上的大量实验表明,X-lens具有卓越的异构相机度量深度精度,在OmniScene-Full上相对于最强基线将AbsRel降低了25.4%,同时使用的参数减少了88.9%,在传统的仅鱼眼和仅针孔设置下也具有竞争力。
英文摘要
We present X-lens, a compact feed-forward model for metric depth estimation from a variable number of calibrated fisheye and pinhole views. To support real-time downstream perception, X-lens is built around a geometry-aware heterogeneous camera formulation with two key components. Learnable calibration tokens provide a coarse alignment between fisheye and pinhole projective spaces, while a Jacobian-parameterized distortion bias injected into cross-attention models local projection changes and promotes cross-camera consistency, enabling robust generalization with only 0.04B parameters and up to 41 FPS. The model predicts dense depth together with a global metric scale, avoiding auxiliary reconstruction targets that increase computation and optimization complexity. To learn such cross-camera generalization at scale and depth, X-lens is trained on multiple public datasets and OmniScene, our newly released large-scale synthetic dataset containing approximately 266K synchronized six-view frames, 1.7M individual images, and 103 indoor and outdoor scenes. Extensive experiments on both real-world and synthetic indoor and outdoor datasets demonstrate superior heterogeneous-camera metric depth accuracy, reducing AbsRel by 25.4\% on OmniScene-Full over the strongest baseline while using 88.9\% fewer parameters, with competitive performance on conventional fisheye-only and pinhole-only settings.
Comments24 pages