发表机构
University of Chinese Academy of Sciences; Zhejiang University; Peking University; Ant Group(中国科学院大学; 浙江大学; 北京大学; 蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究旨在解决高维视觉表示离散化难题,提出超球面量化(HSQ),其离散表示自编码器(dRAE)能解耦语义与特征幅度,防止码本坍缩,实现高保真重建,实验证明性能提升、码本利用率高且训练流程简化。
AI 中文摘要
在这项工作中,我们旨在将高维视觉表示离散化,以弥合与语言模型之间的差距,这是一项艰巨的挑战,因为现有的量化方法存在码本坍缩问题,在保持语义连贯的同时无法扩展。我们将根本原因确定为度量不匹配:标准欧几里得码本目标与表示空间的各向异性几何结构根本不匹配,导致码本嵌入具有高方差幅度尺度和不均匀的角度分布,阻碍了可扩展性。为了解决这个问题,我们提出了超球面量化(HSQ),它通过角度路由将语义内容与特征幅度解耦,防止码分配由尺度而非意义主导。由此产生的离散表示自编码器(dRAE)在保持语义完整性并支持可扩展码本预算的同时实现了高保真重建。广泛的实验表明,随着词汇量扩展到131,072,性能持续提升,码本利用率达100%,简化了训练流程,在理解和生成任务中表现出色。
英文摘要
In this work, we aim to discretize the high-dimensional visual representations to bridge the gap with language models - a non-trivial challenge, as existing quantization methods suffer from codebook collapse, failing to scale while preserving semantic coherence. We identify the root cause as metric mismatch: standard Euclidean codebook objectives are fundamentally misaligned with the anisotropic geometry of representation space, leading to codebook embeddings with high-variance magnitude scales and uneven angular distributions that hinder scalability. To address this, we propose Hyper-Spherical Quantization (HSQ), which decouples semantic content from feature magnitude via angular routing, preventing code assignment from being dominated by scale rather than meaning. The resulting discrete Representation Autoencoder (dRAE) achieves high-fidelity reconstruction while preserving semantic integrity and supporting scalable codebook budget. Extensive experiments demonstrate consistent performance gains as the vocabulary size scales to 131{,}072, along with 100\% codebook utilization, simplified training pipeline, and strong performance across understanding and generation tasks.
CommentsPreprint. Project Page: https://drae-hsq.github.io