发表机构
Carnegie Mellon University; Peking University; Google(卡内基梅隆大学; 北京大学; 谷歌)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
OpWeave提出灵活算子解聚框架,通过分析成本模型和正则性感知规划器优化异构LLM服务,分别降低1.78倍和1.89倍成本。
AI 中文摘要
LLM服务系统日益将推理过程解聚为更细粒度的阶段,近期方法在解码过程中将注意力与FFN或MoE执行分离。这种算子级解聚服务(ODS)可以改善硬件匹配并支持独立扩展,尤其是在异构设备之间。然而,现有系统固定算子边界,且缺乏对解聚何时降低服务成本的统一刻画。我们提出OpWeave,一个面向异构ODS的端到端框架。OpWeave提供了一个分析成本模型,界定了同构和异构ODS相对于共置服务的收益上限。它通过一个正则性感知规划器联合优化算子划分和部署配置,即使对于混合注意力模型也能保持搜索的可处理性。一个基于vLLM的运行时在异构设备组上执行合成的计划,具有灵活的算子阶段。在我们的评估中,相对于最佳可行基线,OpWeave在同构GPU集群上将服务成本降低了最多1.78倍,在异构GPU集群上降低了最多1.89倍,同时满足延迟SLO。
英文摘要
LLM serving systems increasingly disaggregate inference into finer-grained stages, with recent approaches separating attention from FFN or MoE execution during decode. This operator-level disaggregated serving (ODS) can improve hardware matching and enable independent scaling, particularly across heterogeneous devices. However, existing systems fix operator boundaries and lack a unified characterization of when disaggregation reduces serving cost. We present OpWeave, an end-to-end framework for heterogeneous ODS. OpWeave provides an analytical cost model that bounds the gains of homogeneous and heterogeneous ODS over colocated serving. It jointly optimizes operator partitioning and deployment configuration through a regularity-aware planner that keeps the search tractable even for hybrid-attention models. A vLLM-based runtime executes the synthesized plans with flexible operator stages across heterogeneous device groups. In our evaluation, OpWeave reduces serving cost by up to $1.78\times$ on homogeneous and $1.89\times$ on heterogeneous GPU clusters relative to the best feasible baseline, while meeting latency SLOs.
Comments24 pages, 14 figures, including references and appendices