arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

向量、乘积与标量量化的统一率-失真视角

A Unified Rate-Distortion Perspective on Vector, Product, and Scalar Quantization

Xianghong Fang, Wenlong Mou, Yuan Yuan, Dehan Kong, Tim G. J. Rudner

arXiv 2609.02107首次发表:更新:

发表机构

University of Toronto; Boston College(多伦多大学; 波士顿学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对离散视觉分词的VQ、PQ、SQ提出统一率-失真视角,明确重建保真度核心目标,确立量化比较公平条件,还原失真层级并证实现代VQ方法失真最低。

AI 中文摘要

离散视觉分词主要由向量量化(Vector Quantization, VQ)、标量量化(Scalar Quantization, SQ)和乘积量化(Product Quantization, PQ)驱动,但缺乏用于理解量化权衡的统一概念框架。本文针对现代离散视觉分词提出一种统一的率-失真视角,将量化视为有损压缩,通过分词数量和码本大小表征名义固定长度编码率,将量化误差表征为失真。在该框架内,我们解决三个核心问题:第一,从理论和实证上证明,最小化失真而非最大化码本利用率是重建保真度的主要内在目标,这与直通估计器(Straight-Through Estimator, STE)诱导的梯度差异直接相关;第二,为内在量化比较确立两个关键公平条件:控制潜在特征统计量和强制相同编码率;第三,在这些条件下,我们还原现代视觉分词中的VQ-PQ-SQ失真层级,并通过实证表明现代VQ方法实现最低失真。本研究为现代离散视觉分词提供了基础的率-失真重构,解决了量化器评估中的歧义问题,并提供了在固定率约束下分离内在量化有效性的可控框架。

英文摘要

Discrete visual tokenization, predominantly driven by vector, scalar, and product quantization, lacks a unified conceptual framework for understanding quantization tradeoffs. In this paper, we propose a unified rate--distortion perspective on modern discrete visual tokenization. By viewing quantization as lossy compression, we characterize the nominal fixed-length coding rate through token count and codebook size, and quantization error as the distortion. Within this framework, we resolve three central questions. First, we theoretically and empirically show that minimizing distortion, rather than maximizing codebook utilization, is the primary intrinsic objective for reconstruction fidelity, with a direct connection to the STE-induced gradient discrepancy. Second, we establish two critical fairness conditions for intrinsic quantization comparison: controlling latent feature statistics and enforcing identical coding rates. Third, under these conditions, we recover the VQ--PQ--SQ distortion hierarchy in modern visual tokenization and show empirically that modern VQ methods achieve the lowest distortion. This work provides a foundational rate--distortion reframing of modern discrete visual tokenization, resolves ambiguities in quantizer evaluation, and provides a controlled framework for isolating intrinsic quantization effectiveness under fixed-rate constraints.

Comments26 pages, 2 figure, 8 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑