arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

NAMVIS:下一尺度自回归多视图图像合成

NAMVIS: Next-Scale Autoregressive Multi-View Image Synthesis

Ramil Khafizov, Ilya Statsenko, Ruslan Rakhimov, Artem Komarichev, Peter Wonka, Evgeny Burnaev

arXiv 2610.04722首次发表:更新:

发表机构

Applied AI Institute; T-Tech; KAUST(应用人工智能研究所; T-Tech; 阿卜杜拉国王科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

NAMVIS提出无扩散的几何条件化下一尺度自回归框架,通过并行令牌预测实现稀疏视图多视图合成,在多个数据集上超越扩散基线且速度快3倍以上。

AI 中文摘要

稀疏视图的新视图合成是三维内容创作中的一个核心问题,但基于扩散的方法受限于迭代去噪,导致多视图生成在推理时成本高昂。我们提出了NAMVIS,一个无扩散框架,将多视图图像合成重新定义为几何条件化的下一尺度自回归。NAMVIS不是通过重复去噪生成目标视图,而是通过少量从粗到细的尺度步骤预测离散视觉令牌,同时并行采样每个尺度内及跨目标视图的所有令牌。为了将该生成过程锚定到显式相机几何,我们提出了多尺度投影姿态编码,该编码在每一分辨率下将源和目标相机变换注入到目标视图自注意力以及源到目标交叉注意力中。NAMVIS进一步将全局条件与密集的几何感知交叉注意力相结合,使模型能够在保持目标视图一致性的同时保留源视图外观。在Objaverse、GSO和OmniObject3D上,NAMVIS在PSNR、SSIM和LPIPS指标上优于基于扩散的基线,并且在相同评估设置下运行速度比所评估的扩散基线快3倍以上。这些结果表明,几何条件化的下一尺度自回归是稀疏视图多视图合成中一种有前景且高效的扩散替代方案。额外的定性结果、视频和资源可在该https URL获取。

英文摘要

Sparse-view novel view synthesis is a central problem in 3D content creation, but diffusion-based approaches remain limited by iterative denoising, making multi-view generation expensive at inference time. We introduce NAMVIS, a diffusion-free framework that reformulates multi-view image synthesis as geometry-conditioned next-scale autoregression. Instead of generating target views through repeated denoising, NAMVIS predicts discrete visual tokens through a small number of coarse-to-fine scale steps, while sampling all tokens within each scale and across target views in parallel. To anchor this generation process to explicit camera geometry, we propose Multi-scale Projective Pose Encoding, which injects source and target camera transformations into both target-view self-attention and source-to-target cross-attention at every resolution. NAMVIS further combines global conditioning with dense geometry-aware cross-attention, enabling the model to preserve source-view appearance while maintaining target-view consistency. Across Objaverse, GSO, and OmniObject3D, NAMVIS outperforms diffusion-based baselines in PSNR, SSIM, and LPIPS, while running over 3 times faster than the evaluated diffusion baselines under the same evaluation setting. These results suggest that geometry-conditioned next-scale autoregression is a promising and efficient alternative to diffusion for sparse-view multi-view synthesis. Additional qualitative results, videos, and resources are available at https://corl-team.github.io/namvis/

CommentsAccepted at NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑