发表机构
Nanjing University(南京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出将逐光线球形占用轮廓作为统一中间表示,训练判别式解码器与生成式流水线,在多视图三维重建任务中取得良好效果,且可迁移至真实照片场景。
AI 中文摘要
我们研究球形占用轮廓——即从多视图三维高斯重建中提取的逐光线占用概率轮廓P(r)=T(r)∘o(r)——作为从图像进行判别式和生成式三维重建的统一中间表示。在包含999个物体、每个物体有48个转盘视图的Google Scanned Objects子集上,我们训练了:(i) 一个判别式逐光线解码器,它将全局视图平均和逐光线的图像证据注入FiLM条件化的轮廓头,在独立的90个物体测试分割上达到0.035(归一化)的中位软深度误差;(ii) 一个基于轮廓VAE和潜在扩散模型的生成式流水线,支持匹配重建流形的无条件采样,以及通过无分类器引导可量化且可调的每物体解扩散的图像条件多解重建。我们进一步分析预测轮廓的形态:事后锐化和学习到的锐化目标均能在不降低深度的情况下恢复真实轮廓宽度,揭示了L1逐光线损失族中的单调宽度-峰值边界,并推动了形态门的原则性重新定义。在两个DTU场景上的真实照片验证证实了该流水线可迁移至非合成输入。我们的结果表明,逐光线占用轮廓提供了一种紧凑、可学习且可感知不确定性的多视图重建与生成先验之间的接口。
英文摘要
We study spherical occupancy profiles-the ray-wise occupancy probability profiles P(r) = T(r) o(r) distilled from multi-view 3D Gaussian reconstructions-as a unified intermediate representation for both discriminative and generative 3D reconstruction from images. On a 999-object subset of Google Scanned Objects with 48 turntable views each, we train (i) a discriminative per-ray decoder that injects global view-averaged and ray-specific image evidence into a FiLM-conditioned profile head, reaching median soft depth error 0.035 (normalized) on an independent 90-object test split, and (ii) a generative pipeline built on a profile VAE and a latent diffusion model, which supports unconditional sampling that matches the reconstruction manifold and image-conditioned multi-solution reconstruction whose per-object solution spread is quantifiable and tunable via classifier-free guidance. We further analyze the morphology of predicted profiles: post-hoc power sharpening and a learned sharpening target both recover ground-truth profile width without degrading depth, exposing a monotonic width-peak frontier in the L1-per-ray loss family and motivating a principled redefinition of morphology gates. Real-photo validation on two DTU scenes confirms the pipeline transfers to non-synthetic input. Our results suggest that ray-wise occupancy profiles offer a compact, learned, and uncertainty-aware interface between multi-view reconstruction and generative priors.
Comments15 pages, 3 figures, 9 tables. Code and weights: https://github.com/102324988/LSOP_code_release