arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.24436cs.CV

MSVS-VAE:用于高保真3D重建的多尺度锚定向量集

MSVS-VAE: Multi-Scale Anchored VecSet for High-Fidelity 3D Reconstruction

Dehao Hao, Kaiyi Zhang, Tanghui Jia, Xiangjun Gao, Dongyu Yan, Weikai Chen, Zeyu Hu, Lingting Zhu, Yingda Yin, Runze Zhang, Li Yuan, Xin Wang, Long Quan

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对高保真3D重建中VAE的瓶颈问题,提出MSVS-VAE。通过分层点洗牌上采样 densify 锚定向量集潜在特征,用AVS-Conv取代全局交叉注意力,引入多尺度查询解码融合特征,实验表明其性能优于现有方法,解码速度和紧凑性有显著提升。

中文摘要 AI 辅助

高保真3D生成建模越来越依赖潜在扩散范式,其中底层3D VAE的重建质量成为主要瓶颈。现有方法主要遵循两种范式:基于稀疏体素的表示具有强大的重建质量,但会产生大量内存和计算开销;基于集合的表示紧凑且连续,但由于潜在稀疏性和过度全局平滑性,保真度通常较低。我们提出了MSVS-VAE,一种基于层次集合的VAE,在不牺牲紧凑性的情况下弥合了保真度差距。我们的关键思想是通过分层点洗牌上采样逐步 densify 锚定向量集潜在特征,增加用于细粒度几何建模的空间容量。为了从 densified 层次结构中高效解码,我们用AVS-Conv取代全局交叉注意力,AVS-Conv是一种在局部邻域内运行的几何感知局部聚合算子,而不是详尽的潜在集。我们还引入了多尺度查询解码来融合从粗到细的潜在特征,粗尺度提供稳定的全局上下文,细尺度细化局部几何,减少过度局部感受野产生的伪影。在Objaverse、ABO和野外基准上的大量实验表明,MSVS-VAE始终优于先前基于集合和基于体素的VAE,解码速度比先前基于集合的方法快约10倍,紧凑性比基于体素的基线高约10倍。

英文摘要

High-fidelity 3D generative modeling increasingly relies on the latent diffusion paradigm, where the reconstruction quality of the underlying 3D VAE becomes a primary bottleneck. Existing approaches largely follow two paradigms: sparse voxel-based representations achieve strong reconstruction quality but incur significant memory and computational overhead, while set-based representations are compact and continuous yet typically lag in fidelity due to latent sparsity and excessive global smoothness. We propose MSVS-VAE, a hierarchical set-based VAE that closes this fidelity gap without sacrificing compactness. Our key idea is to progressively densify anchored VecSet latents via hierarchical point-shuffle upsampling, increasing spatial capacity for fine-grained geometry modeling. To efficiently decode from the densified hierarchy, we replace global cross-attention with AVS-Conv, a geometry-aware local aggregation operator operating within local neighborhoods rather than the exhaustive latent set. We further introduce multi-scale query decoding to fuse coarse-to-fine latent features, where coarse scales provide stable global context, and fine scales refine localized geometry, reducing artifacts from overly local receptive fields. Extensive experiments on Objaverse, ABO, and in-the-wild benchmarks demonstrate that MSVS-VAE consistently outperforms prior set-based and voxel-based VAEs, delivering approximately 10x faster decoding than prior set-based methods and approximately 10x higher compactness than voxel-based baselines.

发表机构

  • The Hong Kong University of Science and Technology(香港科技大学)
  • Peking University(北京大学)
  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
  • Lightspeed Studios(光速工作室)

机构由 AI 辅助整理,请以论文原文为准。

↑