arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HHR:面向高效大语言模型生成的分层哈希检索

HHR: Hierarchical Hash Retrieval for Efficient LLM Generation

Lianjun Liu, Tiantian Zheng, You Huang, Weiqi Yan, Mingte Qiu, Huazhong Liu, Xiaofeng Zhu, Yunshan Zhong

arXiv 2610.01230首次发表:更新:

发表机构

Hainan University; Xiamen University; Zhejiang Normal University(海南大学; 厦门大学; 浙江师范大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对长上下文推理中哈希检索与注意力相关性不匹配的问题,提出分层哈希检索(HHR)框架,通过几何感知键路由和学习哈希投影提升检索精度,在LongBench上平均分提升1.10,并实现高达3.30倍解码加速。

AI 中文摘要

高效的长上下文推理对于大语言模型(LLMs)至关重要,然而它构成了严重的计算瓶颈。基于哈希的检索通过将查询和键编码为二进制码,并利用汉明距离进行键选择,提供了一种高效的替代方案。然而,这导致了汉明距离与注意力相关性之间的关键不匹配。查询-键logits同时依赖于方向相似性和特征幅度,而哈希二值化丢弃了幅度信息,导致低logit键的误报检索和高logit键的漏报遗漏。为解决这些问题,我们提出了分层哈希检索(HHR),这是一个由粗到细的框架,通过几何感知键路由(GKR)和学习的哈希投影(LHP)逐步提高检索精度。GKR学习一种头级正交变换,以重新分配特征幅度并推导更具判别性的页级logit边界,从而在保留重要候选键的同时有效剪除低logit键。随后,LHP学习一个头级投影空间,使汉明距离与真实的查询-键相关性排序对齐,以实现细粒度检索。通过结合GKR和LHP,HHR抑制了误报并恢复了漏报,显著提高了基于哈希的稀疏注意力的保真度。在多种LLMs和基准上的广泛实验表明,HHR取得了优于现有方法的性能。例如,在LongBench上,HHR将平均得分提高了1.10分,并且在上下文长度为128K时,对于Llama-3.1-8B-Instruct,实现了高达3.30倍的解码加速和2.83倍的端到端加速。代码可在该https URL公开获取。

英文摘要

Efficient long-context inference is essential for large language models (LLMs), yet it poses a severe computational bottleneck. Hash-based retrieval offers an efficient alternative by encoding queries and keys into binary codes and using Hamming distance for key selection. However, this leads to a critical mismatch between Hamming distance and attention relevance. Query-Key logits depend jointly on directional similarity and feature magnitudes, whereas hash binarization discards magnitude information, causing both false-positive retrieval of low-logit keys and false-negative omission of high-logit keys. To address these failures, we propose Hierarchical Hash Retrieval (HHR), a coarse-to-fine framework that progressively improves retrieval accuracy through Geometry-Aware Key Routing (GKR) and Learned Hash Projection (LHP). GKR learns a head-wise orthogonal transformation to redistribute feature magnitudes and derive more discriminative page-level logit bounds, enabling effective pruning of low-logit keys while preserving important candidates. LHP then learns a head-wise projection space that aligns Hamming distance with the true Query-Key relevance ranking for fine-grained retrieval. By combining GKR and LHP, HHR suppresses false positives and recovers false negatives, substantially improving the fidelity of hash-based sparse attention. Extensive experiments across diverse LLMs and benchmarks demonstrate that HHR achieves superior performance over existing methods. For example, on LongBench, HHR improves the average score by 1.10 points and, at a context length of 128K, achieves up to a 3.30x decoding speedup and a 2.83x end-to-end speedup for Llama-3.1-8B-Instruct. The code is publicly available at https://github.com/lianjunl13-sudo/HHR.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑