arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.02201cs.CVcs.AIcs.LG

SILSA:用于保持拓扑的高分辨率3D生成的滑动窗口切片潜变量

SILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generation

Tianjiao Yu, Xinzhuo Li, Yifan Shen, Ying Shen, Kiet A. Nguyen, Adheesh Sunil Juvekar, Ismini Lourentzou

首次发表
浏览论文内容

中文总结 AI 辅助

SILSA提出滑动窗口切片潜变量与切片级拓扑监督,实现单阶段高分辨率3D生成,提升结构保真度并大幅降低计算成本。

中文摘要 AI 辅助

高分辨率3D生成日益依赖于体素潜变量和多阶段流水线,这些流水线首先预测活跃结构,然后合成局部几何。尽管这种方法有效,但它将连续表面碎片化为许多局部令牌,增加了生成成本,并且常常削弱薄或高度连接形状的拓扑一致性。我们引入了SILSA,一个拓扑感知的3D生成框架,使用紧凑的滑动窗口切片潜变量来表示形状。SILSA不生成昂贵的体素令牌,而是沿着三个规范轴使用一组固定的重叠切片,其中每个令牌总结一个局部深度窗口,以保持横截面连续性并支持单阶段整流流生成。一个切片VAE将定向表面样本编码为多轴切片潜变量,并通过稀疏体积解码器重建它们,而一个体积锚点格通过共享的3D工作空间协调定向切片流。为了保持结构正确性,我们引入了切片级拓扑监督,匹配持久性图并对齐相邻切片的Betti转变。实验表明,SILSA在显著降低生成成本的同时提高了结构保真度。与最强基线相比,SILSA将PSNR提高了8.7%,覆盖率提高了5.96个绝对点,Betti误差降低了9.2%,同时使用的令牌比次紧凑基线少70.0%,比稀疏或分层令牌器少超过98%,有效减少了40.4%的训练内存和58.5%的推理时间。定性结果进一步显示了对薄结构、重复组件和长程连通性的改善保留。

英文摘要

High-resolution 3D generation increasingly relies on voxel latents and multi-stage pipelines that first predict active structure and then synthesize local geometry. While effective, this design fragments continuous surfaces into many local tokens, inflates generation cost, and often weakens topological consistency for thin or highly connected shapes. We introduce SILSA, a topology-aware 3D generation framework that represents shapes with compact sliding-window slice latents. Instead of generating expensive voxel tokens, SILSA uses a fixed set of overlapping slices along the three canonical axes, where each token summarizes a local depth window to preserve cross-sectional continuity and support single-stage rectified-flow generation. A Slice VAE encodes oriented surface samples into multi-axis slice latents and reconstructs them with a sparse volumetric decoder, while a Volumetric Anchor Lattice coordinates directional slice streams through a shared 3D workspace. To preserve structural correctness, we introduce slice-level topology supervision that matches persistence diagrams and aligns Betti transitions across neighboring slices. Experiments show that SILSA improves structural fidelity while substantially reducing generation cost. SILSA improves PSNR by $8.7\%$, coverage by $5.96$ absolute points, and Betti error by $9.2\%$ over the strongest baseline, while using $70.0\%$ fewer tokens than the next-most compact baseline and over $98\%$ fewer tokens than sparse or hierarchical tokenizers, effectively reducing training memory by $40.4\%$ and inference time by $58.5\%$. Qualitative results further show improved preservation of thin structures, repeated components, and long-range connectivity.

发表机构

  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑