arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HYPIC: 利用位置无关缓存加速混合注意力LLM服务

HYPIC: Accelerating Hybrid-Attention LLM Serving with Position-Independent Caching

Yifei Liu, Juntong Wu, Yang Liu, Junhao Hu, Minghao Li, Xiaoxu Chen, Weihang Chen

arXiv 2607.01299首次发表:更新:

AI 中文总结

提出Hypic系统,通过段累积转移算子和边界重计算,实现混合注意力LLM的位置无关缓存,显著降低首令牌延迟并提升吞吐量。

AI 中文摘要

在检索增强生成(RAG)和智能体LLM服务中,提示词由独立片段组装成长上下文,使得预填充阶段主导每次请求的计算成本。针对这一成本,并行出现了两个方向:位置无关缓存(PIC)允许对不同请求间共享的非连续片段进行KV重用,而混合注意力模型通过用线性注意力替换大部分全注意力层来降低计算复杂度。然而,它们无法共存:将PIC应用于混合注意力模型会失效,因为每令牌KV缓存重用原语无法迁移到每请求的循环状态。在这项工作中,我们提出了Hypic,这是第一个支持位置无关缓存的混合注意力LLM服务系统。对于线性注意力层,我们识别出段累积转移算子作为缺失的代数原语,并将其与每个片段的零起始结束状态一起缓存,从而实现对独立缓存片段的近乎精确和常数时间的状态组合。对于剩余的全注意力层,现有的PIC方法也失败,因为线性层不暴露每令牌隐藏状态以供选择性重计算。我们表明,最显著的注意力偏差集中在片段边界,因此仅在每个边界重计算一个小接缝窗口就足以恢复跨片段回看。最后,Hypic利用段级自包含性跨实例并行化缓存未命中预填充,将长冷请求(在前缀缓存和先前PIC下都是尾部延迟的主要贡献者)转变为可加速的工作负载。在四种混合注意力模型和五种工作负载上的评估表明,与现有系统相比,Hypic平均将首令牌延迟(TTFT)降低2.45倍,峰值吞吐量提升高达2.0倍,同时精度保持在完全重计算的3.3个百分点以内。

英文摘要

In retrieval-augmented generation and agentic LLM serving, prompts are assembled from independent segments into long contexts, making the prefill stage dominate per-request cost. Two directions have emerged to reduce this cost: position-independent caching (PIC) admits KV reuse for non-contiguous segments shared across requests, while hybrid-attention models cut computation by replacing most full-attention layers with linear attention. However, they cannot coexist: applying existing PIC methods to hybrid-attention models breaks down because per-token KV-cache reuse primitives do not transfer to the per-request recurrent state. We present Hypic, the first system to accelerate hybrid-attention LLM serving with position-independent caching. For linear-attention layers, we identify the segment-cumulative transition operator as the missing algebraic primitive and cache it alongside each segment's zero-start end-state, enabling near-exact and constant-time composition of independently cached segments. For the remaining full-attention layers, existing PIC methods also fail because linear layers do not expose the per-token hidden states needed for selective recomputation. We show that the largest deviations concentrate at segment beginnings and construct a small seam window that propagates hidden states through the hybrid-attention stack to repair cross-segment attention. Finally, Hypic introduces segment parallelism, which exploits PIC's segment-level self-containment to parallelize cache-miss prefill across instances, turning long cold requests into an accelerable workload. Evaluated across four hybrid-attention models and five workloads, Hypic reduces time-to-first-token by $3.25\times$ on average and improves QPS by $1.66\times$ over Prefix Cache, while preserving task quality with a 1.71-point gap from Full Recompute.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑