超越后量化:使用专用哈希令牌的原生哈希学习
Beyond Post-Quantization: Native Hash Learning with a Dedicated HASH Token
浏览论文内容
中文总结 AI 辅助
研究高效图像检索的紧凑表示问题,提出HashViT框架,通过专用哈希令牌及相关设计在Transformer内生成二进制表示,经统一目标优化,实验表明其性能优且保留汉明码效率。
中文摘要 AI 辅助
高效大规模图像检索需要紧凑表示以在快速汉明空间搜索下保持语义相似性。深度哈希有吸引力,但多数现有方法遵循后量化范式,存在特征到代码的差异。为此提出HashViT用于原生哈希令牌学习,介绍专用HASH令牌及相关设计,经统一目标优化,实验验证其性能。
英文摘要
Efficient large-scale image retrieval requires compact representations that preserve semantic similarity under fast Hamming-space search. Deep hashing is appealing, but most existing CNN- and ViT-based methods still follow a post-quantization paradigm, where continuous visual features are first learned and binary codes are then produced by a terminal hash projection or binarization operation. This late code generation creates a feature-to-code discrepancy between the continuously optimized representation space and the discrete Hamming space used for retrieval. To address this limitation, we propose HashViT, a Vision Transformer framework for native hash token learning. Instead of treating hashing as a terminal readout, HashViT introduces a dedicated HASH token that serves as a persistent, hash-oriented retrieval state inside the transformer. The HASH token is structurally decomposed into a Hash Register for direct binary code generation and a Semantic Workspace for preserving auxiliary continuous semantics. To enable effective workspace-to-register interaction, we further design a lightweight Hash Refinement Adapter that progressively refines the Hash Register across transformer layers. As a result, binary-oriented representations are formed through token evolution within the backbone, rather than being abruptly induced by an output-level projection. HashViT is optimized with a unified objective that combines learnable semantic center supervision, class-token similarity distillation, and quantization regularization, encouraging the HASH token to encode semantically structured and compact binary representations. Extensive experiments on three widely used benchmarks demonstrate that HashViT achieves state-of-the-art or highly competitive retrieval performance while preserving the efficiency of compact Hamming codes. Code is available at https://github.com/Xinze919/HashViT.
发表机构
- Institute of Information Engineering, CAS(中国科学院信息工程研究所)
- School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络空间安全学院)
- Department of Applied Mathematics and Statistics, Johns Hopkins University(约翰霍普金斯大学应用数学与统计系)
机构由 AI 辅助整理,请以论文原文为准。