发表机构
Cornell University; Adobe; University of California, Berkeley; Georgia Institute of Technology(康奈尔大学; 奥多比公司; 加州大学伯克利分校; 佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对消费级多视图相机,提出稀疏视图3DGS与4DGS基线方法,验证多相机稀疏采样可显著提升单帧、少帧及casual视频的三维与四维重建质量,尤其适用于动态场景。
AI 中文摘要
许多消费级智能手机、立体相机和光场相机在单次曝光事件中记录多个同步视点。然而,新视图合成管道通常仅使用单目流,依赖相机运动或学习到的先验来获取角覆盖范围。在本文中,我们提出疑问:为何仅使用一个视点?我们分析了传感器受限多视图(其中单个传感器在空间分辨率和角分辨率之间进行权衡)以及曝光受限多视图(其中单个消费设备上的多个传感器同时观测每个事件)。我们引入了一个包含三种类型消费级多视图相机的新数据集,并评估了稀疏视图 3DGS 和 4DGS 基线,以曝光次数和极端视图之间的角度为函数测量重建质量。我们的结果表明,使用多个相机,即使基线较低,也能在单帧、少帧和 casual 视频设置中显著提升重建质量。此外,在固定传感器预算下,尽管空间分辨率较低,但在曝光不足时,角采样可提升重建质量。这种提升在单帧和动态场景中最为显著,因为静止的单目相机缺乏角多样性来恢复场景几何结构和运动。
英文摘要
Many consumer smartphones, stereo cameras, and light field cameras record multiple synchronized viewpoints in a single exposure event. However, novel view synthesis pipelines commonly use only a monocular stream and rely on camera motion or learned priors to obtain angular coverage. In this paper, we ask: why do we use only one viewpoint? We analyze sensor-limited multi-view, where one sensor trades off spatial and angular resolution, and exposure-limited multi-view, where multiple sensors on one commodity device observe each event simultaneously. We introduce a new dataset incorporating three types of commodity multi-view cameras, and evaluate sparse-view 3DGS and 4DGS baselines measuring reconstruction quality as a function of number of exposures and angle between extreme views. Our results demonstrate that using multiple cameras, even with a low baseline, significantly improves reconstruction quality in single-shot, few-shot, and casual video settings. In addition, under a fixed sensor budget, angular sampling improves reconstruction when exposures are scarce despite lower spatial resolution. The gains are most pronounced for single-shot and dynamic scenes, where a stationary monocular camera lacks the angular diversity to recover scene geometry and motion.
CommentsProject page: https://shamus.li/lightfield-gaussian-splatting