FastE:用于LLM嵌入推理的读出触发式令牌压缩
FastE: Readout-Triggered Token Compression for LLM Embedding Inference
浏览论文内容
中文总结 AI 辅助
FastE是一种无需训练的即插即用方法,通过读出触发式令牌压缩,利用固定阈值和注意力分数排序,在保持嵌入质量的同时大幅降低LLM嵌入推理的计算成本。
中文摘要 AI 辅助
在本研究中,我们识别出最终读出LLM嵌入模型中存在深度相关的前缀冗余,这一点在包括Qwen3-Embedding和Qwen3-VL-Embedding在内的代表性骨干网络中尤为显著。我们发现,移除前缀状态在浅层造成的损害远大于在更深层,这表明随着前缀和读出状态在网络中传播,前缀状态变得越来越可压缩。为此,我们提出了FastE,一种无需训练、即插即用的方法。FastE使用批量平均读出-前缀对齐的共享固定阈值作为轻量级在线启发式,以选择何时进行压缩,并根据读出位置分配给前缀状态的注意力分数对前缀状态进行排序,以确定在后续层中保留哪些状态。我们的评估证明了FastE大幅降低计算成本的能力:在NarrativeQA上使用Qwen3-Embedding-0.6B时,它将解码器骨干FLOPs减少了40.11%,同时保留了Full Forward nDCG@10的99.53%。在五个文本嵌入基准、两个骨干规模和三个跨模态检索任务中,质量-效率权衡可通过最大移除比例直接定制,无需重新训练。我们相信FastE为检索、索引、聚类和多模态表示系统中的可扩展嵌入生成提供了实用价值。
英文摘要
In this study, we identify depth-dependent prefix redundancy in final-readout LLM embedding models, notably across representative backbones including Qwen3-Embedding and Qwen3-VL-Embedding. We find that removing prefix states is substantially more damaging in shallow layers than at greater depth, showing that prefix states become increasingly compressible as the prefix and readout states propagate through the network. To this end, we introduce FastE, a training-free, plug-and-play method. FastE uses a shared fixed threshold on batch-mean readout-prefix alignment as a lightweight online heuristic for selecting when compression occurs, and ranks prefix states by the attention scores they receive from the readout position to determine which states are retained in subsequent layers. Our evaluations demonstrate FastE's ability to substantially reduce computational costs: on NarrativeQA with Qwen3-Embedding-0.6B, it reduces decoder-backbone FLOPs by 40.11% while retaining 99.53% of Full Forward nDCG@10. Across five text embedding benchmarks, two backbone scales, and three cross-modal retrieval tasks, the quality-efficiency trade-off is directly customizable through the maximum removal ratio without retraining. We believe FastE offers practical value for scalable embedding generation in retrieval, indexing, clustering, and multimodal representation systems.
发表机构
- Zhejiang University(浙江大学)
- Ant Group(蚂蚁集团)
- Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security(杭州高新区(滨江)区块链与数据安全研究院)
机构由 AI 辅助整理,请以论文原文为准。