发表机构
University of California, San Diego(加州大学圣迭戈分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
QuadTok提出了一种层次化四叉树视觉分词器,动态分配表示能力,节省约10%的标记,并利用树结构因果性实现自回归图像生成,在ImageNet上达到2.08 gFID,支持零样本空间控制生成。
AI 中文摘要
我们提出了QuadTok,一个用于视觉分词和自回归图像生成的新框架。与使用2D网格或1D标记序列的传统方法相比,我们提出了一种层次化的四叉树结构,弥合了2D空间绑定与1D序列级灵活性之间的差距。QuadTok分词器动态地将表示能力分配给视觉复杂的区域,同时将同质区域保留在粗分辨率下。与固定的256标记网格相比,我们的在ImageNet上训练的分词器在ImageNet上节省了约10%的标记,在零样本迁移到COCO数据集时节省了9%,同时保持了相当的重建保真度。此外,树结构引入的自然因果性无缝地实现了自回归图像生成。在生成前提供四叉树拓扑的条件下,我们的947M GPT风格生成模型在ImageNet $256 \ imes 256$基准上达到了2.08的gFID。此外,利用四叉树结构保留的强空间相关性,QuadTok生成器实现了零样本的空间控制图像生成能力。代码:此https URL。
英文摘要
We introduce QuadTok, a novel framework for visual tokenization and autoregressive image generation. Compared to traditional approaches using 2D grids or 1D token sequences, we propose a hierarchical quadtree structure, bridging the gap between 2D spatial binding and 1D sequence-level flexibility. The QuadTok tokenizer dynamically allocates representational capacity to visually intricate areas while leaving homogeneous regions at a coarse resolution. Compared with a fixed 256-token grid, our ImageNet-trained tokenizer saves approximately 10% of tokens on ImageNet and 9% when transferred zero-shot to the COCO dataset, while maintaining comparable reconstruction fidelity. Furthermore, the natural causality introduced by the tree structure seamlessly enables autoregressive image generation. Conditioned on a quadtree topology supplied before generation, our 947M GPT-style generative model achieves a 2.08 gFID on the ImageNet $256 \times 256$ benchmark. Additionally, leveraging the strong spatial correlation preserved by the quadtree structure, the QuadTok generator enables zero-shot spatially controlled image generation capabilities. Code: https://github.com/myc634/QuadTok.