用于矢量量化的分布匹配:一个统一的理论和实证框架
Distributional Matching for Vector Quantization: A Unified Theoretical and Empirical Framework
- University of Toronto(多伦多大学)
- The Hong Kong University of Science and Technology(香港科技大学)
- Boston College(波士顿学院)
- Lehigh University(里海大学)
- Southern University of Science and Technology(南方科技大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对现有矢量量化方法训练不稳定和码本崩溃问题,提出分布匹配框架,通过对齐特征和码向量分布缓解上述问题,经理论分析和实验验证,基于瓦瑟斯坦距离目标实例化该框架,在视觉tokenization基准上表现有效且鲁棒。
AI中文摘要:
现代视觉表征学习和自回归模型的有效性严重依赖矢量量化(VQ),它使用可学习码本离散化连续特征表示。尽管广泛使用,但现有VQ方法常因直通估计器导致的梯度失配和码向量利用不足而存在训练不稳定和码本崩溃问题。本文表明这两个问题可追溯到特征向量和码向量分布的根本不匹配,导致表示效率低下和信息损失。基于此,提出分布匹配框架,引入理想VQ行为的原则标准,经理论分析和实证评估表明对齐特征和码向量分布可缓解训练不稳定和码本崩溃。使用基于瓦瑟斯坦距离的目标在温和高斯近似下实例化该框架,还表明基于最大均值差异的非参数替代方法性能相当。在视觉tokenization基准上的大量实验支持了该方法的有效性和鲁棒性。
英文摘要:
The effectiveness of modern visual representation learning and autoregressive models critically depends on vector quantization (VQ), which discretizes continuous feature representations using a learnable codebook. Despite its widespread use, existing VQ methods often suffer from training instability and codebook collapse, arising from gradient mismatch induced by the straight-through estimator and the under-utilization of code vectors. In this work, we show that both issues can be traced to a fundamental mismatch between the distributions of feature vectors and code vectors, leading to inefficient representation and information loss. Building on this observation, we propose a distributional matching framework for vector quantization. We introduce principled criteria for desirable VQ behavior and demonstrate through theoretical analysis and empirical evaluation that aligning feature and code vector distributions provides a unifying mechanism for mitigating training instability and codebook collapse. We instantiate this framework using a Wasserstein-based objective with an efficient closed-form under a mild Gaussian approximation, and further show that a nonparametric alternative based on maximum mean discrepancy yields comparable performance. Extensive experiments on visual tokenization benchmarks support the effectiveness and robustness of the proposed approach.