arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.09110cs.CVcs.GR

适用于视图一致的2D到3D生成的视图自适应渲染器

View-Adaptive Renderer for View-Consistent 2D-to-3D Generation

U-Chae Jun, Jaeeun Ko, Jiwoo Kang

首次发表
浏览论文内容

中文总结 AI 辅助

针对单目2D转3D生成的视图不一致与计算负担问题,提出视图自适应神经渲染框架,结合自注意力融合模块,在无扩散SDS监督下实现高效高保真3D重建,性能接近当前最优。

中文摘要 AI 辅助

从单张图像重建3D形状仍是计算机视觉领域一个基础但极具挑战性的问题。传统的单目3D生成流程通常先从单张输入图像合成多个视图,再应用基于神经辐射场(NeRF)的重建方法。然而,固有的投影歧义性往往会在生成的视点间产生视觉不连续性,导致重建的3D模型存在误差。当前的解决方案要么会带来显著的额外计算负担,要么无法充分解决合成视图间的实际不一致问题。为应对这些局限,我们提出了一种新颖的视图自适应神经渲染框架,即便在给定部分不一致的多视图输入时,也能实现稳健的3D重建。我们的方法引入了视图自适应神经渲染器,可独立校正依赖视点的误差,同时共享全局特征主干以保持结构一致性。此外,我们提出了一种自注意力融合模块,能自适应整合多视图信息,确保几何一致性,且无需过度依赖间接正则化或计算密集型方法。通过大量实验,我们证明所提方法可持续提升3D重建保真度。重要的是,我们的方法无需基于扩散的SDS监督,主要依靠光度渲染损失与轻量注意力正则化器,即可达到接近当前最优的性能。这种精度与效率间的平衡,使得所提框架在实际应用中极具实用性。

英文摘要

Reconstructing 3D shapes from a single image remains a fundamental yet challenging problem in computer vision. Traditional monocular 3D generation pipelines typically synthesize multiple views from a single input image before applying Neural Radiance Field (NeRF)-based reconstruction. However, inherent projective ambiguities often produce visual discontinuities across generated viewpoints, leading to inaccuracies in reconstructed 3D models. Current solutions either incur significant additional computational burdens or fail to adequately resolve practical inconsistencies between synthesized views. To address these limitations, we propose a novel viewpoint-adaptive neural rendering framework that enables robust 3D reconstruction even when given partially inconsistent multi-view inputs. Our approach introduces view-adaptive neural renderers that independently correct viewpoint-dependent errors while simultaneously sharing a global feature backbone to preserve structural coherence. Furthermore, we propose a self-attention fusion module that adaptively integrates multi-view information, ensuring geometric consistency without relying heavily on indirect regularizations or computationally intensive methods. Through extensive experiments, we demonstrate that our method consistently improves 3D reconstruction fidelity. Importantly, our approach achieves near state-of-the-art performance without diffusion-based SDS supervision, relying primarily on photometric rendering loss with lightweight attention regularizers. This balance between accuracy and efficiency makes the proposed framework highly practical for real-world applications.

补充信息

↑