发表机构
China Mobile Communications Company Limited Research Institute; The University of Tokyo(中国移动通信有限公司研究院; 东京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MEOW提出一种前馈方法,仅凭图像单次重建混合透视、鱼眼和全景相机的三维场景,无需标定或相机类型,通过数据引擎自适应,零样本迁移至真实数据并取得领先精度。
AI 中文摘要
真实世界的数据采集是异构的:透视相机、鱼眼相机和360°全景图像可以在同一个重建任务中共存,然而大多数前馈三维重建模型假设输入为透视图像且格式统一。近期一些能处理多种相机类型的模型,要么需要获知每个视图的相机类型,要么一次仅重建一对图像。目前尚无单次前馈方法能够仅凭图像重建包含完整全景图的混合相机元组。我们提出MEOW,一个前馈系统,它能在单次前馈中,仅凭图像联合重建度量点图和相机位姿,输入为一个混合了透视、鱼眼和全景图像的N视图元组:无需提供任何视图的标定、畸变参数、相机类型标签或位姿。我们的核心设计理念是将异构相机重建视为数据自适应问题,而非架构重新设计。MEOW保留了一个预训练的透视骨干网络,并完全通过一个程序化数据引擎学习异构相机,该引擎在连续的相机模型流形上渲染每个场景,具有精确的光线和深度,并为每个相机采样的训练元组验证共视性。仅使用合成元组训练,MEOW可零样本迁移至真实采集:在异构2D3DS元组上,它达到79.9 mAA@30,而Wid3R在给定每个视图相机类型的情况下为53.8;在我们激光扫描的混合相机基准上,它对每个四视图混合元组实现了79.4 AUC@30的配准。数据引擎、基准和完整评估流程将公开发布。
英文摘要
Real-world capture is heterogeneous: perspective, fisheye, and $360^\circ$ panoramic images can coexist within a single reconstruction task, yet most feed-forward 3D reconstruction models assume perspective imagery and a uniform input representation. Recent models handling several camera types are either informed of the camera type for each view or reconstruct one image pair at a time. No single-pass method reconstructs mixed-camera tuples containing full panoramas from images alone. We present MEOW, a feed-forward system that jointly reconstructs metric pointmaps and camera poses from one N-view tuple mixing perspective, fisheye and full-panorama images, in a single forward pass from images alone: no calibration, distortion parameters, camera-type labels or poses are supplied for any view. Our guiding design philosophy is to treat heterogeneous-camera reconstruction as a data-adaptation problem rather than an architectural redesign. MEOW retains a perspective-pretrained backbone and learns heterogeneous cameras entirely from a procedural data engine, which renders each scene across a continuous manifold of camera models with exact rays and depth, and certifies covisibility for every camera-sampled training tuple. Trained on synthetic tuples only, MEOW transfers zero-shot to real captures: on heterogeneous 2D3DS tuples it achieves 79.9 mAA@30 against 53.8 for Wid3R given the camera type of every view; on our laser-scanned mixed-camera benchmark it registers every four-view mixed tuple with 79.4 AUC@30. The data engine, benchmark, and complete evaluation pipeline will be released.
Comments24 pages, 9 figures, 14 tables