发表机构
Huazhong University of Science and Technology; Meshy AI; Technical University of Munich(华中科技大学; Meshy AI; 慕尼黑工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MeshOctave通过二进空间分辨率定义尺度,采用确定性坍缩和分裂重连操作,结合尺度条件离散扩散模型,实现并行、动态长度的网格生成,显著提升几何保真度和拓扑有效性。
AI 中文摘要
生成具有显式拓扑的紧凑、艺术家风格网格通常依赖于自回归模型,这些模型会产生高昂的逐令牌顺序成本,或依赖于启发式连接解码器的连续流模型。下一代尺度生成范式通过支持尺度内并行令牌预测以及从全局结构到局部拓扑的从粗到细细化,提供了一种引人注目的替代方案;然而,现有方法通过渐进式网格简化来推导分层尺度,并按顺序反转它们。这消除了尺度内并行性,并使生成步骤随面数线性增长。在本文中,我们提出MeshOctave,它转而通过二进空间网格分辨率来定义尺度,将粗化视为一种确定性坍缩,合并共享体素单元的顶点并继承连接性。其逆操作,即分裂与重连,确定每个粗面实例化哪些八分体子顶点,并使用离散结构令牌解析局部连接。这些逐面操作不需要序列化,每个尺度转换被建模为一个无序集合,增加一位坐标精度,自然支持动态长度网格和自适应分辨率细化。我们构建了一个尺度条件的掩码均匀离散扩散模型,以从分辨率坍缩层级中学习分裂与重连操作。MeshOctave在几何保真度和拓扑有效性方面以显著优势超越强基线,同时支持自适应分辨率细化并自然扩展到网格细分任务。
英文摘要
Generating compact, artist-style meshes with explicit topology typically relies on autoregressive models which incur prohibitive sequential per-token costs, or continuous flow models that depend on heuristic connectivity decoders. Next-scale generation paradigms offer a compelling alternative by enabling parallel intra-scale token prediction and coarse-to-fine refinement from global structure to local topology; yet, existing methods derive hierarchical scales via progressive mesh simplification and invert them sequentially. This eliminates intra-scale parallelism and scales generation steps linearly with face count. In this paper, we propose MeshOctave, which instead defines scale through dyadic spatial grid resolutions, framing coarsening as a deterministic collapse that merges vertices sharing a voxel cell and inherits connectivity. Its inverse operation, split-and-rewire, determines which octant sub-vertices are instantiated for each coarse face and resolves local connectivity using discrete structural tokens. These per-face operations require no serialization, each scale transition is modeled as an unordered set that adds one bit of coordinate precision, naturally supporting dynamic-length meshes and adaptive resolution refinement. We construct a scale-conditioned masked-uniform discrete diffusion model to learn split-and-rewire operation from resolution collapse hierarchies. MeshOctave outperforms strong baselines in geometric fidelity and topological validity by a non-trivial margin, while supporting adaptive resolution refinement and extending naturally to mesh subdivision tasks.