arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MeshCarve:紧凑潜空间中的流匹配工艺网格生成

MeshCarve: Artisan Mesh Generation with Flow Matching in Compact Latent Spaces

Xiyu Wang, Ruocheng Wu, Yufei Wang, Zhihao Li, Lanqing Guo, Bihan Wen

arXiv 2610.09723首次发表:更新:

发表机构

Nanyang Technological University; SparAI Inc.; University of Texas at Austin(南洋理工大学; SparAI 公司; 德克萨斯大学奥斯汀分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MeshCarve提出一种在紧凑潜空间中分别生成顶点与边的流匹配方法,通过分层稀疏Transformer和顶点-链接编码缩短令牌序列,在Objaverse上超越现有自回归与流匹配方法,并泛化至Toys4K。

AI 中文摘要

先前的工艺网格生成工作主要采用自回归方式预测面片令牌,这导致推理速度缓慢。近期方法转而通过流匹配对由变分自编码器(VAE)构建的连续潜变量进行建模,但当几何与拓扑被联合编码时,重建质量显著下降,且当潜空间被压缩时问题更为严重。我们提出MeshCarve,一种完全在紧凑潜空间中生成的流匹配方法,分别生成顶点位置与边连接,从而规避了联合紧凑潜空间的难点。为缩短令牌序列,我们提出了一种分层稀疏Transformer骨干,实例化为VertexVAE和EdgeVAE。两种VAE均不编码表面体素上的场,而是锚定于其潜空间中的离散顶点,这大幅缩短了令牌序列长度,而我们的空间感知压缩在不牺牲重建质量的前提下进一步缩短了序列。VertexVAE直接编码顶点占用。对于连接性,我们提出了顶点-链接编码,将顶点之间的任意连接转换为固定长度的逐顶点连续嵌入,并忠实地恢复复杂的艺术拓扑。MeshCarve将这些VAE与锚点生成器结合,并在缩短的令牌序列上进行流匹配。在Objaverse上,它相较于最先进的自回归和流匹配方法展现出优势,并泛化至Toys4K。据我们所知,它是首批每个生成阶段均在空间压缩潜空间中运行的工艺网格生成方法之一,其令牌序列仅为先前最压缩的自回归和流匹配方法的一小部分。

英文摘要

Prior artisan mesh generation works largely predict face tokens autoregressively, which makes inference slow. Recent methods instead flow match continuous latents built by Variational AutoEncoders (VAEs), but reconstruction quality drops significantly when geometry and topology are jointly encoded, and further when the latent space is compressed. We present MeshCarve, a flow matching method that generates entirely in compact latent spaces, generating vertex positions and edge connections separately and sidestepping the difficulty of a joint compact latent. To shorten the token sequence, we propose a hierarchical sparse transformer backbone, instantiated as VertexVAE and EdgeVAE. Instead of encoding fields over the surface voxels, both VAEs anchor on discrete vertices in their latent spaces, which drastically reduces the token sequence length, and our spatial-aware compression shortens it further without costing reconstruction. VertexVAE directly encodes vertex occupancy. For connectivity, we propose vertex-link encoding, which turns arbitrary connectivity between vertices into fixed-length continuous per-vertex embeddings and recovers complex artistic topology faithfully. MeshCarve combines these VAEs with an anchor generator and flow matches on the shortened token sequences. It shows advantages over state-of-the-art autoregressive and flow matching methods on Objaverse and generalizes to Toys4K. To the best of our knowledge, it is among the first artisan mesh generation methods whose every generative stage runs in a spatially compressed latent, with a token sequence only a fraction of the most compressed previous autoregressive and flow matching works.

Comments18 pages, 6 figures, 9 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑