MSVS-VAE:用于高保真3D重建的多尺度锚定向量集
MSVS-VAE: Multi-Scale Anchored VecSet for High-Fidelity 3D Reconstruction
浏览论文内容
中文总结 AI 辅助
研究针对高保真3D重建中VAE的瓶颈问题,提出MSVS-VAE。通过分层点洗牌上采样 densify 锚定向量集潜在特征,用AVS-Conv取代全局交叉注意力,引入多尺度查询解码融合特征,实验表明其性能优于现有方法,解码速度和紧凑性有显著提升。
中文摘要 AI 辅助
高保真3D生成建模越来越依赖潜在扩散范式,其中底层3D VAE的重建质量成为主要瓶颈。现有方法主要遵循两种范式:基于稀疏体素的表示具有强大的重建质量,但会产生大量内存和计算开销;基于集合的表示紧凑且连续,但由于潜在稀疏性和过度全局平滑性,保真度通常较低。我们提出了MSVS-VAE,一种基于层次集合的VAE,在不牺牲紧凑性的情况下弥合了保真度差距。我们的关键思想是通过分层点洗牌上采样逐步 densify 锚定向量集潜在特征,增加用于细粒度几何建模的空间容量。为了从 densified 层次结构中高效解码,我们用AVS-Conv取代全局交叉注意力,AVS-Conv是一种在局部邻域内运行的几何感知局部聚合算子,而不是详尽的潜在集。我们还引入了多尺度查询解码来融合从粗到细的潜在特征,粗尺度提供稳定的全局上下文,细尺度细化局部几何,减少过度局部感受野产生的伪影。在Objaverse、ABO和野外基准上的大量实验表明,MSVS-VAE始终优于先前基于集合和基于体素的VAE,解码速度比先前基于集合的方法快约10倍,紧凑性比基于体素的基线高约10倍。
英文摘要
High-fidelity 3D generative modeling increasingly relies on the latent diffusion paradigm, where the reconstruction quality of the underlying 3D VAE becomes a primary bottleneck. Existing approaches largely follow two paradigms: sparse voxel-based representations achieve strong reconstruction quality but incur significant memory and computational overhead, while set-based representations are compact and continuous yet typically lag in fidelity due to latent sparsity and excessive global smoothness. We propose MSVS-VAE, a hierarchical set-based VAE that closes this fidelity gap without sacrificing compactness. Our key idea is to progressively densify anchored VecSet latents via hierarchical point-shuffle upsampling, increasing spatial capacity for fine-grained geometry modeling. To efficiently decode from the densified hierarchy, we replace global cross-attention with AVS-Conv, a geometry-aware local aggregation operator operating within local neighborhoods rather than the exhaustive latent set. We further introduce multi-scale query decoding to fuse coarse-to-fine latent features, where coarse scales provide stable global context, and fine scales refine localized geometry, reducing artifacts from overly local receptive fields. Extensive experiments on Objaverse, ABO, and in-the-wild benchmarks demonstrate that MSVS-VAE consistently outperforms prior set-based and voxel-based VAEs, delivering approximately 10x faster decoding than prior set-based methods and approximately 10x higher compactness than voxel-based baselines.
发表机构
- The Hong Kong University of Science and Technology(香港科技大学)
- Peking University(北京大学)
- The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
- Lightspeed Studios(光速工作室)
机构由 AI 辅助整理,请以论文原文为准。