arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ZipTok3D:采用紧凑 token 前缀的高保真 3D 分词

ZipTok3D: High-Fidelity 3D Tokenization with Compact Token Prefixes

Mingda Lin, Weijie Wang, Zeyu Zhang, Bowen Cui, Yefei He, Haoyu Zhao, Yuanyu He, Donny Y. Chen, Feng Chen, Bohan Zhuang

arXiv 2609.01740首次发表:更新:

发表机构

Zhejiang University; Monash University; University of Adelaide(浙江大学; 莫纳什大学; 阿德莱德大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有3D分词器在极低token预算下重建性能下降的问题,本文提出ZipTok3D,通过嵌套dropout训练使前导token承载关键几何信息,结合迭代解码,仅用1或4个token就达到基线重建质量,大幅缩短token序列长度。

AI 中文摘要

紧凑的 token 序列对于高效的 3D 生成至关重要。然而,现有的 3D 分词器通常将潜在表示组织为空间区域或固定大小的全局 token 集合,两者在压缩到极低 token 预算时都会出现严重的重建性能下降。本文提出了 ZipTok3D,一种专为从极短 token 序列实现高保真重建而设计的 3D 分词器,其核心思路是将物体几何结构组织为逐步提供信息的全局 token 前缀,并通过迭代解码展开这些紧凑表示。具体而言,嵌套 dropout 在训练时会在编码后随机截断潜在序列,要求每个保留的前缀都能重建完整物体,从而让前导 token 优先承载关键几何信息;解码器则重复应用参数共享的 Transformer 模块,无需单独的生成采样阶段,即可从每个前缀恢复细粒度几何结构。在相同 token 维度下,ZipTok3D 在 ShapeNet 上仅用 1 个 token、在 TRELLIS 上仅用 4 个 token,就达到了 32-token COD-VAE 基线的重建质量,对应 token 序列长度分别缩短了 32 倍和 8 倍。

英文摘要

Compact token sequences are essential for efficient 3D generation. However, existing 3D tokenizers typically organize latent representations either over spatial regions or as fixed-size sets of global tokens, both suffering sharp reconstruction degradation when compressed to extremely low token budgets. In this paper, we present ZipTok3D, a 3D tokenizer designed for high-fidelity reconstruction from extremely short token sequences. Its key idea is to organize object geometry into progressively informative global-token prefixes and unfold these compact representations through iterative decoding. Specifically, nested dropout randomly truncates the latent sequence after encoding during training and requires each retained prefix to reconstruct the complete object, thereby prioritizing essential geometric information in the leading tokens. The decoder then repeatedly applies a parameter-shared Transformer block to recover fine-grained geometry from each prefix without a separate generative sampling stage. With the same token dimension, ZipTok3D achieves reconstruction quality comparable to the 32-token COD-VAE baseline using only one token on ShapeNet and four on TRELLIS, yielding $32\times$ and $8\times$ shorter token sequences, respectively.

Comments25 pages, 9 figures, 6 tables, including appendix

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑