超越全局隐变量:面向可扩展3D建模的基于块的稀疏网格变分自编码器
Beyond Global Latents: Chunk-Based Sparse Grid VAE for Scalable 3D Modeling
浏览论文内容
中文总结 AI 辅助
该研究针对稀疏体素网格内存随分辨率快速增长的问题,提出基于块的ChunkVAE,通过局部算子和互补数据算子实现可扩展3D建模,在多基准测试中性能优于或相当基线,且内存与推理效率更优。
中文摘要 AI 辅助
稀疏体素网格保留了详细3D重建所需的空间结构,但其内存会随分辨率随活跃表面单元的增加而快速增长。我们提出ChunkVAE,这是一种围绕局部块而非全局隐变量体积组织的稀疏网格变分自编码器。局部学习算子允许独立选择编码器和解码器分区,并使推理块大小与训练时不同。两种互补数据算子使这种灵活性成为可能:平衡二值对象划分可分配活跃单元同时限制重复重叠,而S曲线加权拼接在组装全局隐变量或重建时衰减不可靠的边界特征。在三个对象基准测试中,ChunkVAE在$512^3$到$1536^3$的分辨率下与强大基线相当或更优;更小的块降低了峰值分配内存并缩短了每块计算时间,实现更快的并行推理。稳定的拼接隐变量和改进的图像转3D指标表明,局部压缩可在缩放几何结构的同时保留下游所需的全局接口。
英文摘要
Sparse voxel grids preserve the spatial structure needed for detailed 3D reconstruction, but their memory still grows rapidly with resolution as active surface cells increase. We introduce ChunkVAE, a sparse grid variational autoencoder organized around local chunks rather than a global latent volume. Local learned operators permit independently chosen encoder and decoder partitions and allow inference chunk sizes to differ from training. Two complementary data operators make this flexibility practical: Balanced Binary Object Partitioning distributes active cells while limiting replicated overlap, while S-Curve weighted stitching attenuates unreliable boundary features when assembling a global latent or reconstruction. Across three object benchmarks, ChunkVAE is competitive with or better than strong baselines from $512^3$ to $1536^3$; smaller chunks lower peak allocated memory and shorten per-chunk compute, enabling faster parallel inference. Stable stitched latents and improved image to 3D metrics indicate that local compression can scale geometry while retaining the global interface required downstream.
发表机构
- Hong Kong University of Science and Technology(香港科技大学)
- Tencent Hunyuan(腾讯混元)
- The Chinese University of Hong Kong(香港中文大学)
机构由 AI 辅助整理,请以论文原文为准。