发表机构
Nanjing University; Ant Group; Shanghai Jiao Tong University; Peking University(南京大学; 蚂蚁集团; 上海交通大学; 北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Bole是内核-运行时协同设计方案,通过优化树推测解码的线性注意力树验证与瞬态内存管理,显著提升混合注意力LLM的离线解码吞吐量与在线智能体工作负载的延迟性能。
AI 中文摘要
混合注意力大语言模型将全注意力与循环线性注意力结合以降低长上下文推理成本,但其自回归解码仍受内存限制。树推测解码是极具吸引力的加速路径,但现有树推测系统围绕全注意力模型的键值缓存设计,在混合模型中会逐分支遍历循环层并为每个提议节点实例化完整状态,导致验证延迟和瞬态内存随树规模与批次规模增长而急剧恶化。我们提出Bole,一种内核-运行时协同设计,可实现混合注意力LLM的高效树推测解码。Bole将线性注意力循环转换为树结构化闭式形式,通过资源高效的GPU内核实现,并行验证所有提议节点,使线性注意力树验证速度提升3.4至7.7倍;它无损地将推测状态更新编码为token级因子,仅重建采样后选中的状态,使瞬态状态内存减少82至99倍,释放GPU容量用于键值缓存。Bole集成到广泛部署的生产级LLM服务引擎SGLang中,将高效状态管理与针对完整混合前向传播校准的全批次验证预算相结合。在四种模型、两个GPU平台及多样化数据集上,Bole的离线解码吞吐量最高可达自回归解码的4.72倍,是最强树推测基线的2.03倍;在在线智能体工作负载下,其端到端首token延迟(TTFT)和每token输出时间(TPOT)较最强树推测基线分别降低最高67.6%和49.9%。
英文摘要
Hybrid-attention large language models combine full attention with recurrent linear attention to reduce long-context inference costs, yet their autoregressive decoding remains memory-bound. Tree speculative decoding offers an attractive acceleration path, but existing tree-speculation systems are designed around the key--value caches of full-attention models. On hybrid models, they traverse recurrent layers branch by branch and materialize a full state for every proposal node, causing verification latency and transient memory to scale poorly with tree and batch sizes. We present Bole, a kernel--runtime co-design that enables efficient tree speculation for hybrid-attention LLMs. Bole transforms the linear-attention recurrence into a tree-structured closed form and realizes it with a resource-efficient GPU kernel, verifying all proposal nodes in parallel and accelerating linear-attention tree verification by 3.4--7.7$\times$. It losslessly encodes speculative state updates as token-level factors and reconstructs only the state selected after sampling, reducing transient state memory by 82--99$\times$ and freeing GPU capacity for KV caches. Its integration into SGLang, a widely deployed production LLM serving engine, couples efficient state management with a batch-wide verification budget calibrated to the complete hybrid forward. Across four models, two GPU platforms, and diverse datasets, Bole delivers up to $4.72\times$ the offline decode throughput of autoregressive decoding and up to $2.03\times$ that of the strongest tree-speculative baseline. Under online agent workloads, it reduces TTFT and TPOT by up to $67.6%$ and $49.9%$, respectively, over the strongest tree-speculative baseline.
Comments14 pages, 12 figures, 7 tables