发表机构
Fudan University; Alibaba Group(复旦大学; 阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出AffineTok,通过语义仿射一致性(SAC)引导视觉分词器训练,引入M_SAC作为代理指标,在ImageNet 256上实现了gFID的显著降低,达到无引导1.21、有引导1.10的SOTA性能。
AI 中文摘要
视觉分词器日益在隐空间中注入语义监督,以简化下游扩散模型的应用,但如何组织这些语义以促进去噪过程的问题仍未得到充分探索。本文定义了语义恢复目标:去噪过程应从带噪隐变量中恢复干净图像的语义内容,而一个良好的分词器应使这一过程更易实现。现有方法训练投影器直接从带噪隐变量预测语义,我们认为这种方式预测的是干净图像语义的平均值,而真正需要对齐的是平均干净隐变量的语义。更重要的是,我们证明语义恢复误差可正交分解为直接从带噪隐变量进行最优语义预测的误差,以及这两种预测之间的误差,因此我们将它们的一致性确定为缺失的要求,称之为语义仿射一致性(Semantic Affine Consistency, SAC)。为检验这一被忽视的要求是否与下游生成质量密切相关,我们引入了M_SAC,这是一种用于SAC的分词器侧代理指标。在评估的分词器和扩散模型规模范围内,M_SAC与生成质量高度相关,与SiT-XL的gFID达到0.960的皮尔逊相关系数,从而为SAC引导的分词器训练提供了动机。随后我们引入AffineTok,它通过两个互补的仅训练组件来促进SAC:全局语义协调令牌(Global Semantic Coordination Token, GSCT)协调干净隐变量的语义组织,保持语义平均的意义;后验均值语义对齐(Posterior-Mean Semantic Alignment, PMSA)从带噪输入预测后验均值隐变量并对其语义进行监督。在ImageNet 256上,与基线相比,AffineTok在20个训练轮次时将gFID降低了26%,且随着训练的推进,在无无分类器引导的情况下达到了1.21的新SOTA gFID,在有引导的情况下达到了1.10的新SOTA gFID。
英文摘要
Visual tokenizers increasingly inject semantic supervision into latent spaces to make downstream diffusion easier. Yet how these semantics should be organized to facilitate denoising remains underexplored. In this paper, we define the semantic recovery objective: the denoising process should recover the semantic content of the clean image from noisy latent, and a good tokenizer should make it easier. Existing approaches train a projector to predict the semantics directly from the noisy latent. We argue that this predicts the average of clean-image semantics, whereas what really needs to be aligned is the semantics of averaged clean latents. More importantly, we demonstrate that the semantic recovery error orthogonally decomposes into the error of the optimal semantic prediction directly from the noisy latent and the error between these two predictions. We therefore identify their consistency as the missing requirement and call it Semantic Affine Consistency (SAC). To examine whether this overlooked requirement is closely related to downstream generation, we introduce M_SAC, a tokenizer-side proxy for SAC. Across the evaluated tokenizers and diffusion model scales, M_SAC closely tracks generation quality, reaching a Pearson correlation of 0.960 with SiT-XL gFID, thereby motivating SAC-guided tokenizer training. We then introduce AffineTok, which promotes SAC through two complementary, training-only components. Global Semantic Coordination Token (GSCT) coordinates the semantic organization of clean latents, keeping semantic averaging meaningful, while Posterior-Mean Semantic Alignment (PMSA) predicts posterior-mean latents from noisy inputs and supervises their semantics. On ImageNet 256, compared with the baseline, AffineTok reduces gFID by 26% at 20 epochs and, with continued training, achieves a new state-of-the-art gFID of 1.21 without classifier-free guidance and 1.10 with guidance.
CommentsProject page: https://michaelyu781.github.io/AffineTok-site/