AI 中文总结
该研究针对大语言模型推理时的输出嵌入内存瓶颈,提出用基于HNSW的向量索引替代稠密词汇投影,在保持生成质量的同时,使Gemma 3 270M的端到端解码吞吐量提升最高达82%。
AI 中文摘要
大型输出嵌入矩阵在自回归解码过程中会造成显著的内存带宽瓶颈,尤其对于拥有多语言大词汇表的紧凑大语言模型(LLM)而言更为突出。我们将“输出投影后接top-k token选择”重新表述为针对token嵌入的最大内积搜索,并用基于HNSW的向量索引替代稠密词汇投影。由此得到的输出头仅检索少量高分token候选集,可通过将检索到的logits分散到稀疏全词汇张量中,集成到现有解码流程中。在CPU推理环境下,针对Gemma 3、Llama 3.2和Qwen 3模型,我们的方法大幅加快了输出投影速度,使Gemma 3 270M模型的端到端批量大小为1的解码吞吐量提升高达82%,同时在AlpacaEval评估下保持了生成质量。这些结果表明,近似检索是延迟敏感型小批量解码场景中稠密输出投影的实用替代方案。
英文摘要
Large output embedding matrices create a significant memory bandwidth bottleneck during autoregressive decoding, especially for compact LLMs with large multilingual vocabularies. We reformulate the output projection followed by top-k token selection as a maximum inner product search over token embeddings and replace the dense vocabulary projection with an HNSW-based vector index. The resulting output head retrieves only a small candidate set of high-scoring tokens and can be integrated into existing decoding pipelines by scattering retrieved logits into a sparse full-vocabulary tensor. On CPU inference with Gemma 3, Llama 3.2, and Qwen 3 models, our method substantially accelerates the output projection and improves end-to-end batch-size-one decoding throughput by up to 82% for Gemma 3 270M, while preserving generation quality under AlpacaEval evaluation. These results suggest approximate retrieval is a practical alternative to dense output projections in latency-sensitive small-batch decoding.
CommentsICML 2026 - AdaptFM Workshop