arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

P2Voxel:用于3D网格令牌化的金字塔枢轴体素化方法

P2Voxel: Pyramid Pivot Voxelization for 3D Mesh Tokenization

Zhenhong Sun, Haozhe Liu, Yifu Wang, Xibin Song, Senbo Wang, Huadong Mo, Daoyi Dong, Hongdong Li, Pan Ji

arXiv 2608.07549首次发表:更新:

发表机构

Australian National University; Vertex Lab; University of New South Wales; University of Technology Sydney(澳大利亚国立大学; 顶点实验室; 新南威尔士大学; 悉尼科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出P2Voxel框架,通过三项关键创新将网格令牌化为紧凑结构化的金字塔枢轴令牌,实现高效网格重建,适用于下游3D任务。

AI 中文摘要

三角形网格能提供明确且精确的表面几何信息,但其不规则的拓扑连通性使得3D网格令牌化成为一个几何采样问题:如何将几何证据采样并组织成紧凑、结构化且可学习的令牌。不同于以场为中心的体积采样和边相交表面采样,本文将网格令牌化重新定义为「局部表面证据采样」,即识别每个活跃体素内足以实现确定性表面恢复的最小几何证据。为此,本文提出P2Voxel,一种用于紧凑且感知重建的网格令牌化的金字塔枢轴体素化框架。P2Voxel基于三项关键创新:在「局部平面性」假设下,枢轴体素化用一个表面枢轴和一个方向符号表示每个活跃体素,提供可诱导确定性重建所需角值的最小局部证据;在「空间复杂性」假设下,金字塔枢轴体素化利用真实表面的空间非均匀性,在几何复杂区域分配更精细的枢轴令牌,同时保持平滑区域的粗粒度和紧凑性;在「块可重建性」假设下,金字塔VAE在局部可重建的枢轴块上学习紧凑的多分辨率潜在代码,无需将整个高分辨率体素化形状建模为密集全局场。这些设计共同将网格转换为紧凑、结构化且可学习的金字塔枢轴令牌,为下游3D任务实现高效的网格重建。

英文摘要

Triangle meshes provide explicit and accurate surface geometry, yet their irregular topology connectivity makes 3D mesh tokenization a geometric sampling problem: how to sample and organize geometric evidence into compact, structured and learnable tokens. Beyond field-centric volumetric sampling and edge-intersection surface sampling, we retarget mesh tokenization as \textit{local surface evidence sampling}: identifying the minimal geometric evidence inside each active voxel that is sufficient for deterministic surface recovery. To this end, we introduce \textbf{P2Voxel}, a pyramid pivot voxelization framework for compact and reconstruction-aware mesh tokenization. P2Voxel is built on three key innovations. Under the \textit{Local Planarity} assumption, Pivot Voxelization represents each active voxel with a surface pivot and an orientation sign, providing minimal local evidence that can induce the corner values required for deterministic reconstruction. Under the \textit{Spatial Complexity} assumption, Pyramid Pivot Voxelization exploits the spatial non-uniformity of real surfaces by allocating finer pivot tokens to geometrically complex regions while keeping smooth regions coarse and compact. Under the \textit{Block Reconstructability} assumption, a Pyramid VAE learns compact multi-resolution latent codes over locally reconstructable pivot blocks, avoiding the need to model the entire high-resolution voxelized shape as a dense global field. Together, these designs convert meshes into compact, structured, and learnable pyramid pivot tokens, enabling efficient mesh reconstruction for downstream 3D tasks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑