发表机构
North China Electric Power University; University of Science and Technology of China; Stanford University(华北电力大学; 中国科学技术大学; 斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
NesTok提出嵌套自对齐框架,通过跨长度训练提升一维变长分词器的重建与生成性能,在ImageNet上取得rFID 0.98和gFID 1.46的领先结果。
AI 中文摘要
一维(1D)变长视觉分词器通过改变分词数量实现自适应压缩,使得下游自回归(AR)模型能够利用单个分词器灵活地在生成质量与计算成本之间进行权衡。然而,现有基于嵌套丢弃的方法往往未能充分利用分词器的表示能力,导致图像重建和生成性能均欠佳。在本工作中,我们引入了NesTok,一种专为动态视觉分词器设计的嵌套自对齐框架。NesTok引入了跨长度训练,该训练在多个分词长度上联合优化重建,同时利用全长序列引导较短序列,使较短分词序列能够接近全长序列的重建质量。在ImageNet上,NesTok相较于标准训练有显著提升,并取得了0.98的rFID分数。在下游图像生成中,它在ImageNet 256×256上取得了现有变长自回归图像生成方法中最先进的gFID分数1.46。代码将在此https URL提供。
英文摘要
One-dimensional (1D) variable-length visual tokenizers enable adaptive compression by varying the number of tokens, allowing downstream autoregressive (AR) models to flexibly trade off generation quality against computational cost using a single tokenizer. However, existing approaches based on nested dropout often fail to fully exploit the representational capacity of the tokenizer, resulting in suboptimal performance in both image reconstruction and generation. In this work, we introduce NesTok, a nested self-alignment framework tailored to dynamic visual tokenizers. NesTok introduces cross-length training, which jointly optimizes reconstruction across token lengths while using the full-length sequence to guide shorter counterparts, enabling shorter token sequences to approach the reconstruction quality of full-length sequences. On ImageNet, NesTok improves substantially over standard training and achieves an rFID score of 0.98. On downstream image generation, it achieves the state-of-the-art gFID score of 1.46 on ImageNet 256$\times$256 among existing variable-length autoregressive image generation methods. Code will be available at https://github.com/Jiawei804/NesTok.
CommentsComputer Vision, Autoregressive Model