arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03641cs.CV

用于高效渐进式图像压缩的树结构矢量量化

Tree-VQ: Progressive Image Compression from Pretrained Vector Quantizers

Mingming Ma, Xinkun Wang, Tianyi Xu, Qingyu Luo, Fu Li, Yi Niu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出Tree-VQ框架,通过树结构矢量量化实现渐进式图像压缩,其性能效率优于对比方法,参数更少、延迟更低且感知压缩效果最佳。

中文摘要 AI 辅助

基于矢量量化的图像压缩已取得优异的率失真性能,但大多数方法仍为每个目标比特率生成独立的压缩表示。这种可变速率特性使单个模型能在多个速率下运行,但不一定生成前缀可解码且可通过附加比特进行细化的渐进式比特流。我们提出Tree-VQ,一种用于学习图像压缩的渐进式树结构矢量量化框架。Tree-VQ将离散码字组织为分层二叉树,并用路由的根到叶路径表示每个潜在令牌。关键在于,该路径的每个前缀都对应一个有效的量化表示,因此浅层内部节点作为粗略重建码,更深节点提供连续细化。这使得压缩图像可从早期前缀解码,并在接收到更多分支符号时逐步改进,而非为不同目标比特率重新编码。为使该结构适用于压缩,我们引入前缀兼容树熵模型,仅使用因果可用的解码上下文对渐进式延续决策和路由分支细化进行编码。我们进一步采用感知率的细化调度,以在给定前缀预算下决定哪些空间块应接收额外树比特,以及分层前缀监督以确保内部节点在低速率下可直接解码。实验表明,Tree-VQ实现了优异的性能效率权衡,与对比方法相比,以少得多的参数和更低延迟提供了最佳感知压缩结果。

英文摘要

Progressive image compression requires a single embedded representation whose received prefixes can be decoded without re-encoding the source. Modern vector-quantized (VQ) image models provide strong discrete endpoint representations, but conventional flat codeword indices do not define meaningful intermediate states for a neural decoder. We present Tree-VQ, a post-hoc conversion of a pretrained flat VQ tokenizer into a fine-grained, arbitrary-prefix progressive representation while preserving its encoder assignments and every learned leaf vector. The key idea is to organize the original codebook into a balanced binary hierarchy, associate explicit representations with internal nodes, and transmit branch decisions in depth-major order. Consequently, once the image header is available, every payload prefix uniquely specifies a valid latent state: each additional branch bit refines exactly one token, and transmission can therefore be truncated at essentially any payload position rather than only at a small number of stage boundaries. We further adapt one shared decoder on the complete-depth and mixed-depth latent states encountered under such arbitrary truncation, making these densely spaced prefixes useful for reconstruction rather than merely syntactically decodable. On Kodak, Tree-VQ achieves a DISTS-based BD-rate saving of 52.1% relative to ProGIC, while exposing thousands of valid arbitrary-prefix operating points from a single embedded bitstream.

发表机构

  • School of Artificial Intelligence(人工智能学院)
  • Xidian University(西安电子科技大学)

机构由 AI 辅助整理,请以论文原文为准。

↑