arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26774cs.CV

StableVQ:稳定向量量化分词器训练的实用指南

StableVQ: Practical Guidelines for Stable Vector-Quantized Tokenizer Training

Bao Tang, Jiahao Guo, Haoxiang Cao, Wenyu Liu, Changqian Yu, Kun Gai, Xinggang Wang

首次发表
浏览论文内容

中文总结 AI 辅助

StableVQ通过解耦编码器-解码器与码本的学习目标,提出动态STE、区域VQ损失和解耦调度,在不增加参数的情况下提升VQ分词器训练的稳定性与码本利用率。

中文摘要 AI 辅助

向量量化(VQ)是离散视觉分词器的基础,这些分词器为现代自回归和掩码图像生成模型提供动力。尽管最近的共享投影码本方法已大幅提升了码本利用率,但训练稳定性仍然是一个关键且未被充分探索的挑战。我们认为,根本原因在于编码器-解码器与码本训练的纠缠:由于任一模块都无法在孤立状态下可靠地履行自身职责,系统只能在两个子系统恰好合作时才能运行——而这种脆弱条件恰好在训练压力最大时崩溃。我们提出StableVQ,重新审视每个模块的适当学习目标,并解决当每个模块独立履行自身角色时出现的问题。具体而言,(1)动态STE(Dynamic STE)纠正了编码器学习目标中的不稳定性,使其能够在离散正则化下稳健地优化重建空间,即使码本利用率较低。(2)区域VQ损失(Region VQ Loss)重新构想码本的学习目标,使其能够独立保证完全跟踪编码器输出分布,而不依赖编码器振荡来驱动激活。(3)解耦调度(Decoupled Schedule)认识到编码器-解码器与码本的不同职责需要不同的优化动态,并为每个模块分配独立的学习率调度,以确保稳健的系统级行为。基于共享投影码本构建,StableVQ轻量且不引入可学习参数。在ImageNet上的实验表明,在不同码本大小和初始化设置下,训练稳定性、码本利用率和重建质量均得到一致提升。

英文摘要

Vector Quantization (VQ) is fundamental to discrete visual tokenizers that power modern autoregressive and masked image generation models. While recent shared-projection codebook methods have substantially advanced codebook utilization, training stability remains a critical and underexplored challenge. We argue that the root cause lies in the entanglement of the Encoder--Decoder and Codebook training: because neither module can reliably fulfill its own responsibility in isolation, the system can only function when the two subsystems happen to cooperate---a fragile condition that breaks down precisely when training is most stressed. We propose StableVQ, which revisits the proper learning objective of each module and resolves the problems that arise when each is trained to fulfill its own role independently. Concretely, (1) Dynamic STE corrects the instability in the Encoder's learning objective, enabling it to robustly optimize the reconstruction space under discrete regularization even when codebook utilization is low. (2) Region VQ Loss reconceives the Codebook's learning objective so that it can independently guarantee full tracking of the encoder output distribution, without relying on encoder oscillations to drive activation. (3) Decoupled Schedule recognizes that the distinct responsibilities of the Encoder--Decoder and the Codebook demand distinct optimization dynamics, and assigns each an independent learning rate schedule to ensure robust system-level behavior. Built on top of shared-projection codebooks, StableVQ is lightweight and introduces no learnable parameters. Experiments on ImageNet demonstrate consistent improvements in training stability, codebook utilization, and reconstruction quality across diverse codebook sizes and initialization settings.

发表机构

  • Huazhong University of Science and Technology(华中科技大学)
  • KlingAI Research(可灵AI研究院)
  • South China Normal University(华南师范大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑