谱局部敏感哈希:通过Krylov投影局部敏感哈希实现次二次方的提示压缩
Spectral-LSH: Sub-Quadratic Prompt Compression via Krylov-Projected Locality-Sensitive Hashing
浏览论文内容
中文总结 AI 辅助
研究针对长提示推理成本高的问题,提出谱局部敏感哈希方法,通过Krylov子空间方法等近似注意力核算子,应用SimHash分组聚合令牌,实验揭示压缩率相变及不同压缩率下方法优势,自适应后端可兼顾不同情况。
中文摘要 AI 辅助
长提示推理成本高昂,因为预填充注意力随序列长度呈二次方扩展。我们提出谱局部敏感哈希(Spectral-LSH),这是一种无需训练的提示压缩方法,在提示进入语言模型之前运行。Spectral-LSH使用Krylov子空间方法和随机特征近似隐式注意力核算子的主要成分,避免显式的$O(N^2)$注意力核实例化。然后在得到的注意力特征空间中应用SimHash对相似令牌进行分组,并将它们聚合为具有因果位置分配的宏令牌。我们在C4上对Mistral-7B-Instruct-v0.3、Qwen2.5-7B-Instruct和Qwen2.5-14B-Instruct进行评估。实验揭示了压缩率相变。低于$\rho = 4 \times$,局部令牌冗余足够低,轻量级分块通常提供最佳延迟-质量权衡。高于$\rho = 8 \times$,谱路径保留了分块丢失的质量。在$\rho = 16 \times$时,Qwen2.5-7B(自适应)将困惑度从353.409降至196.963,而Qwen2.5-14B(自适应)将其从9.533降至3.427。在包含类似JSON、类似代码和类似表格输入的小型长上下文结构化压力测试中,局部局部敏感哈希在$8 \times$时也比分块在每个指标上有所改进。自适应后端通过在低压缩时使用分块路径和在高压缩时使用谱聚类来捕捉两种情况,尽管分块在总延迟方面仍然是最快的后端。
英文摘要
Long-prompt inference remains expensive because prefill attention scales quadratically with sequence length. We propose Spectral-LSH, a training-free prompt compression method that operates before the prompt enters the language model. Spectral-LSH approximates the dominant components of an implicit attention-kernel operator using a Krylov subspace method together with random features, avoiding explicit $O(N^2)$ attention-kernel materialization. It then applies SimHash in the resulting attention eigenspace to group similar tokens and aggregate them into macro-tokens with causal positional assignments. We evaluate Mistral-7B-Instruct-v0.3, Qwen2.5-7B-Instruct, and Qwen2.5-14B-Instruct on C4. Our experiments reveal a compression-ratio phase transition. Below $ρ= 4 \times$, local token redundancy is low enough that lightweight chunking typically provides the best latency--quality trade-off. Above $ρ= 8 \times$, the spectral path preserves quality that chunking loses. At $ρ= 16 \times$, Qwen2.5-7B (adaptive) reduces the PPL ratio from 353.409 to 196.963, while Qwen2.5-14B (adaptive) reduces it from 9.533 to 3.427. On a small long-context structured stress test containing JSON-like, code-like, and table-like inputs, local LSH also improves every metric over chunking at $8 \times$. The adaptive backend captures both regimes by using the chunk path at low compression and spectral clustering at high compression, although chunking remains the fastest backend in total latency.
发表机构
- Islamic Azad University(伊斯兰阿扎德大学)
- Iran University of Science and Technology(伊朗科技大学)
- Meta
机构由 AI 辅助整理,请以论文原文为准。