从透视图像到鱼眼图像的深度估计与开放词汇分割
From Perspective to Fisheye Depth Estimation and Open-Vocabulary Segmentation
浏览论文内容
中文总结 AI 辅助
该研究提出了与架构、任务无关的Distortion Extenders(DEX),通过自监督对齐损失将视觉基础模型泛化到鱼眼相机,在深度估计和开放词汇分割任务中优于基线,还可用于相机校准。
中文摘要 AI 辅助
视觉基础模型能够对三维(3D)场景进行高保真度估计并实现跨场景泛化,其实验成功源于在大规模透视图像数据集上的训练。然而,当将这些模型迁移到鱼眼相机拍摄的宽视场(FoV)图像时,由于图像像素上的径向畸变引发的协变量偏移,模型会输出错误结果。我们提出了一种将视觉基础模型泛化到鱼眼相机的方法,核心是一组可学习参数,称为畸变扩展器(Distortion Extenders, DEX),该参数用于建模鱼眼畸变系数以及鱼眼图像与透视图像在潜在空间中的分布偏移。通过最小化自监督对齐损失,DEX将鱼眼图像的潜在嵌入转换为与透视图像潜在嵌入相似的形式,从而恢复高保真度估计。DEX与架构和任务无关:我们在基于卷积和Transformer的架构上,针对单目深度估计和开放词汇分割任务验证了DEX,在室内和室外鱼眼数据集上,DEX均持续优于基线方法。作为附带成果,DEX的激活值还可被解码为畸变系数,以支持相机校准。代码可访问:this https URL。
英文摘要
Vision foundation models are capable of generalizing across 3-dimensional (3D) scenes with high-fidelity estimates; their empirical success can be attributed to training on large-scale datasets of perspective images. However, when transferred to wide field-of-view (FoV) images, such as those captured by fisheye cameras, they return erroneous outputs due to a covariate shift stemming from the radial distortion on the image pixels. We propose a method to generalize vision foundation models to fisheye cameras. The crux of our method lies in a set of learnable parameters, termed Distortion Extenders (DEX), that model the fisheye distortion coefficients and the distributional shift between fisheye and perspective images encoded in the latent space. By minimizing a self-supervised alignment loss, DEX transforms the latent embeddings of fisheye images to resemble those of perspective images to recover high-fidelity estimates. DEX is architecture- and task-agnostic: We demonstrate DEX on monocular depth estimation and open-vocabulary segmentation for convolution- and Transformer-based architectures, where we consistently improve over baselines across indoor and outdoor fisheye datasets. As a byproduct, the activations of DEX can also be decoded to distortion coefficients to support camera calibration. Code available at: https://github.com/Suchisrit/DEX.
发表机构
- Yale University(耶鲁大学)
机构由 AI 辅助整理,请以论文原文为准。