arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MIVIFI:弥合透视域与鱼眼域以训练多视图鱼眼图像生成模型

MIVIFI: Bridging Perspective and Fisheye Domains for Training Multi-View Fisheye Image Generation Models

Matthias Neuwirth-Trapp, Begüm Altunbas, Jiayi Wang, Yan Xia, Maarten Bieshaar, Xinyu Huang, Daniel Cremers

arXiv 2608.23140首次发表:更新:

发表机构

ETH Zurich; Bosch Research; Technical University of Munich(苏黎世联邦理工学院; 博世研究中心; 慕尼黑工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对多视图鱼眼图像生成问题,提出SyntheOcc-FE与MIVIFI两种方法,其中MIVIFI利用跨域学习缓解鱼眼数据稀缺问题,实现高保真的多视图鱼眼图像生成。

AI 中文摘要

实现360°覆盖对自动驾驶车辆的视觉感知系统至关重要,鱼眼相机仅需两个传感器即可实现全环绕覆盖,是一种高性价比的解决方案。然而,现有的多视图鱼眼数据集有限,合成罕见的边缘案例通常需要计算成本高昂的3D模拟,这阻碍了模型的训练。尽管生成模型在标准透视图像领域已取得显著成功,但其在广角畸变场景中的应用尚未得到探索。在本研究中,我们正式提出了基于体积语义表示的多视图鱼眼图像生成这一新问题,并给出两种不同的方法。我们首先提出SyntheOcc-FE,该方法将SyntheOcc架构适配到鱼眼数据中,虽有效,但受限于鱼眼数据集的稀缺性,限制了其泛化能力。为克服这些局限,我们提出第二种方法MIVIFI(多视图鱼眼),其利用带有等角矩形投影的跨域学习,通过结合KITTI-360鱼眼图像与nuScenes多视图标准图像弥合数据集域之间的差距,使我们的方法能够实现场景内容的高保真操作。该框架支持对语义占据输入进行结构修改,以引入或移除特定对象,并能渲染有限鱼眼数据集中不存在的各类气象条件与光照场景。定量与定性实验表明,我们的方法可实现鲁棒的照片级真实感多视图鱼眼图像生成,并凸显了跨域策略在应对数据稀缺问题上的特定优势。

英文摘要

Achieving 360° coverage is critical for the visual perception systems of autonomous vehicles. Fisheye cameras offer a cost-effective solution by enabling full surround coverage with as few as two sensors. However, existing multi-view fisheye datasets are limited, and synthesizing rare corner cases typically requires computationally expensive 3D simulations, hindering the training. While generative models have achieved significant success in standard perspective imagery, their application to wide-angle distortion remains unexplored. In this work, we formally introduce the novel problem of multi-view fisheye image generation conditioned on volumetric semantic representations and present two distinct methods. We first propose SyntheOcc-FE, which adapts the SyntheOcc architecture to fisheye data. While effective, this method is constrained by the scarcity of fisheye datasets, which limits its generalization. To overcome these limitations, we propose our second method, MIVIFI (multi-view fisheye), which leverages cross-domain learning with Equirectangular Projections. By bridging the gap between dataset domains using KITTI-360 fisheye images alongside nuScenes multi-view standard images, our approach enables high-fidelity manipulation of scene content. This framework enables the structural modification of semantic occupancy inputs to introduce or eliminate specific actors and facilitates the rendering of diverse meteorological conditions and illumination scenarios absent in the limited fisheye datasets. Quantitative and qualitative experiments demonstrate that our methods achieve robust photorealistic multi-view fisheye image generation and highlight the specific advantages of our cross-domain strategy for handling data scarcity.

CommentsAccepted at the IEEE International Conference on Intelligent Transportation Systems (ITSC) 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑