发表机构
Meshy AI(Meshy AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Meshy T2是基于流匹配的网格生成框架,通过顶点集网格VAE和粗到细流匹配级联实现快速图像到网格生成,速度比自回归基线快一个数量级以上,兼具高几何保真度与面数控制、多部件支持能力。
AI 中文摘要
多边形网格是现代3D管线的标准表面表示,生成具有艺术家风格拓扑的高质量网格对电影、游戏和交互式3D应用至关重要。主流方法将网格序列化为标记序列并自回归解码,推理速度慢且易受误差累积影响,不适合交互式资产创建。我们提出Meshy T2,一个基于流匹配的快速原生网格生成框架,其核心是顶点集网格VAE,它将每个顶点编码为一个连续潜标记,并一次性解码顶点、边连接性和面缠绕顺序,无需顶点量化或焊接即可保留高精度几何和艺术家创作的拓扑结构。生成过程分为两个流匹配模型的粗到细级联:图像条件体素流首先将整体形状勾勒为粗略占用支架,然后网格流在该支架中填充每个顶点的潜标记,该过程以图像、支架和请求的顶点预算为条件。此设计提供三项实用能力:基于并行流合成的交互式生成速度;通过请求的顶点预算实现的有效面数控制;以及原生支持多部件资产,其组件直接从生成的连接性中产生。在实验中,Meshy T2实现了最先进的几何保真度,端到端图像到网格生成的中位数时间为6秒,比自回归基线快一个数量级以上。代码和权重将在此https URL获取。
英文摘要
Polygonal meshes are the standard surface representation of modern 3D pipelines, and generating high-quality meshes with artist-style topology is essential for film, gaming, and interactive 3D applications. Mainstream approaches serialize a mesh into a token sequence and decode it autoregressively, which is slow at inference and sensitive to error accumulation, making them impractical for interactive asset creation. We present Meshy T2, a fast native mesh generation framework built on flow matching. At its core is a vertex-set mesh VAE that encodes a mesh into one continuous latent token per vertex and decodes vertices, edge connectivity, and face winding order in a single pass, preserving high-precision geometry and artist-authored topology without vertex quantization or welding. Generation proceeds as a coarse-to-fine cascade of two flow-matching models: an image-conditioned voxel flow first sketches the overall shape as a coarse occupancy scaffold, and a mesh flow then populates the scaffold with per-vertex latent tokens, conditioned on the image, the scaffold, and a requested vertex budget. This design delivers three practical capabilities: interactive generation speed through parallel flow-based synthesis; effective face-count control through the requested vertex budget; and native support for multi-part assets, whose components emerge directly from the generated connectivity. In our experiments, Meshy T2 achieves state-of-the-art geometric fidelity and completes end-to-end image-to-mesh generation within a median of 6 seconds, over an order of magnitude faster than autoregressive baselines. Code and weights will be available at https://github.com/meshy-dev/meshy-t2.