arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VQ-Transplant:用于预训练视觉分词器的高效VQ模块集成

VQ-Transplant: Efficient VQ-Module Integration for Pre-trained Visual Tokenizers

Xianghong Fang, Yuan Yuan, Dehan Kong, Tim G. J. Rudner

arXiv 2607.19575首次发表:更新:

发表机构

University of Toronto; Boston College(多伦多大学; 波士顿学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对训练VQ模块资源需求大的问题,提出VQ-Transplant框架,可将新VQ模块集成到预训练分词器,保留编码器-解码器参数,引入轻量级解码器适应策略,降低训练成本95%,在工业级模型上获高重建保真度,推动量化研究发展。

AI 中文摘要

向量量化(VQ)是现代离散视觉分词的基础。然而,为基于VQ的先进模型训练量化模块需要大量计算资源,这在实际中几乎阻碍了资源受限下新型前沿VQ技术的发展。为解决此限制,我们提出VQ-Transplant,一个简单框架,通过替换预训练分词器的原生VQ模块,实现新VQ模块的即插即用集成。关键是,移植过程保留所有编码器-解码器参数,修改量化方法时无需进行昂贵的端到端重新训练。为减轻解码器-量化不匹配,我们引入轻量级解码器适应策略(在ImageNet-1k上仅训练5个epoch)来使特征先验与新量化空间对齐。在实证评估中,VQ-Transplant能让VAR等工业级模型获得接近最优的重建保真度,同时将训练成本降低95%。VQ-Transplant通过实现资源高效集成新型VQ技术并匹配工业级重建性能,使量化研究更普及。

英文摘要

Vector Quantization (VQ) underpins modern discrete visual tokenization. However, training quantization modules for state-of-the-art VQ-based models requires significant computational resources which, in practice, all but prevents the development of novel, cutting-edge VQ techniques under resource constraints. To address this limitation, we propose {\bf VQ-Transplant}, a simple framework that enables plug-and-play integration of new VQ modules into frozen, pre-trained tokenizers by replacing their native VQ modules. Crucially, the proposed transplantation process preserves all encoder-decoder parameters, obviating the need for costly end-to-end retraining when modifying the quantization method. To mitigate decoder-quantization mismatch, we introduce a lightweight decoder adaptation strategy (trained for only 5 epochs on ImageNet-1k) to align feature priors with the new quantization space. In our empirical evaluation, we find that VQ-Transplant allows obtaining near state-of-the-art reconstruction fidelity for industry-level models like VAR while reducing the training cost by 95\%. VQ-Transplant democratizes quantization research by enabling resource-efficient integration of novel VQ techniques while matching industry-level reconstruction performance.

Comments20 pages, 9 figures and 16 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑