arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GLIDE:用于高效大语言模型推理的引导式分层混合注意力机制

GLIDE: Guided Layerwise Hybrid Attention for Efficient LLM Inference

Vimal William, Ravi Tandon, Jyotikrishna Dass

arXiv 2607.24788首次发表:更新:

AI 中文总结

研究针对大语言模型推理时KV缓存瓶颈问题,提出GLIDE引导式分层混合注意力机制,通过分层自适应平衡线性循环与softmax窗口,非均匀压缩softmax占用空间,实现性能-效率权衡,降低长上下文生成延迟。

AI 中文摘要

随着大语言模型处理的上下文越来越长,解码过程中键值(KV)缓存的内存I/O和计算开销成为主要的吞吐量瓶颈。为解决此问题,我们提出了GLIDE,一种引导式分层混合注意力机制,它将滑动窗口softmax注意力与线性循环聚合策略性地集成。GLIDE受分层异质性启发,早期层对去除softmax敏感,深层则有冗余且能容忍线性替代。它引入分层自适应机制,各层平衡高效线性循环与可变大小softmax窗口。与均匀混合方法不同,GLIDE非均匀压缩模型中的softmax占用空间,减少总KV缓存I/O,同时在关键处保留表达能力。实验评估表明,GLIDE在性能-效率权衡方面表现出色,降低长上下文生成的端到端延迟且不影响质量。

英文摘要

As Large Language Models scale to increasingly long contexts, the memory I/O and computational overhead of the Key-Value (KV) cache during decoding emerges as the primary throughput bottleneck. To address this, we propose GLIDE, a Guided Layerwise Hybrid Attention that strategically integrates sliding-window softmax attention with linear recurrent aggregation. GLIDE is motivated by layer-wise heterogeneity: early layers exhibit high sensitivity to softmax removal, while deeper layers demonstrate redundancy and tolerate aggressive replacement by linear alternatives. Leveraging this insight, GLIDE introduces a layer-wise adaptive mechanism wherein each layer balances an efficient linear recurrence with a variable-sized softmax window. Unlike uniform hybrid approaches, GLIDE non-uniformly compresses the softmax footprint across the model, reducing aggregate KV cache I/O while preserving expressive power where most vital. Empirical evaluations demonstrate the GLIDE achieves superior performance-efficiency tradeoffs, reducing end-to-end latency for long-context generation without compromising quality.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑