OneLA:在生成式推荐中将线性注意力解码扩展到大规模束搜索
OneLA: Scaling Linear-Attention Decoding to Large Beams in Generative Recommendation
浏览论文内容
中文总结 AI 辅助
OneLA提出线性注意力解码框架,利用共享提示和发散记录表示束状态,减少内存与数据移动,实现1.54-2.46倍端到端解码加速。
中文摘要 AI 辅助
生成式推荐(GR)依赖于大束宽解码来生成数百个候选项目,这为循环线性注意力带来了新的扩展挑战。现有的线性注意力服务系统要么为每个束物化完整的循环状态,要么重复重放共享的历史记录,从而产生大量的内存和流量开销。为解决此问题,我们提出了OneLA,一个线性注意力解码框架,它利用了GR工作负载的共享提示和短发散后缀。具体来说,OneLA使用一个单一的共享提示派生状态和紧凑的仅追加的发散转换记录来表示所有束状态。利用这种表示,OneLA仅计算每个解码步骤所需的状态信息,而无需为每个束重建完整的循环状态。此外,OneLA使用轻量级祖先索引来跟踪构成每个束历史的转换记录,使得束可以在不移动或复制现有记录的情况下进行更新。一个融合的GPU内核进一步在束之间重用共享状态。我们的分析表明,OneLA实现了1.54-2.46倍的端到端解码加速,同时大幅减少了循环状态内存使用和数据移动。
英文摘要
Generative recommendation (GR) relies on large-beam decoding to generate hundreds of candidate items, creating a new scaling challenge for recurrent linear attention. Existing linear attention serving systems either materialize a full recurrent state for every beam or repeatedly replay shared history, incurring substantial memory and traffic overhead. To address this, we present OneLA, a linear-attention decoding framework that exploits the shared prompt and short divergent suffixes of GR workloads. Specifically, OneLA represents all beam states using a single shared prompt-derived state and compact, append-only records of their divergent transitions. Using this representation, OneLA computes only the state information required at each decoding step, without reconstructing a full recurrent state for every beam. Furthermore, OneLA uses a lightweight ancestry index to track the transition records that make up each beam's history, allowing beams to be updated without moving or copying existing records. A fused GPU kernel further reuses the shared state across beams. Our analysis shows that OneLA achieves 1.54-2.46x end-to-end decode speedups while substantially reducing recurrent-state memory use and data movement.
发表机构
- The University of Hong Kong(香港大学)
- Kuaishou Technology(快手科技)
- Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。