AI 中文总结
针对大语言模型服务系统,HeraSys通过细粒度编排消除工作流间计算冗余,引入负载感知联合调度策略,结合资源倾斜机制等,有效减轻尾部延迟并提高吞吐量,实验证明其能显著降低延迟、提升服务吞吐量。
AI 中文摘要
大语言模型的扩散使服务系统从处理孤立请求转向编排高并发、多租户的智能工作流。然而,现有解决方案通常优先考虑工作流内优化,很大程度上忽视了工作流间优化的巨大潜力。本文提出HeraSys,一个旨在优化并发工作流端到端性能的大语言模型服务系统。通过细粒度编排,HeraSys通过结构节点合并和重用消除工作流间计算冗余。此外,HeraSys引入负载感知联合调度策略,通过评估查询间和查询内优先级动态管理执行顺序。通过将资源倾斜机制与自适应批处理和流水线分解相结合,HeraSys在保持低平均延迟的同时有效减轻尾部延迟,从而大幅提高系统吞吐量。大量实验表明,在严格的延迟保证下,HeraSys将P99延迟降低了2.17倍,并将服务吞吐量提高了1.85倍。
英文摘要
The proliferation of Large Language Models (LLMs) has shifted serving systems from processing isolated requests to orchestrating high-concurrency, multi-tenant agentic workflows. However, existing solutions typically prioritize intra-workflow optimization, largely neglecting the significant potential for inter-workflow optimization. In this paper, we propose HeraSys, an LLM serving system designed to optimize the end-to-end performance of concurrent workflows. Through fine-grained orchestration, HeraSys eliminates cross-workflow computational redundancy via structural node merging and reuse. Furthermore, HeraSys introduces a load-aware joint scheduling policy that dynamically manages execution order by evaluating both inter- and intra-query priorities. By integrating a resource skewing mechanism with adaptive batching and pipeline decomposition, HeraSys effectively mitigates tail latency while maintaining low average latency, thereby substantially improving system throughput. Extensive experiments demonstrate that HeraSys reduces P99 latency by up to 2.17$\times$ and increases serving throughput by up to 1.85$\times$ under strict latency guarantees.
Commentsto be published in ICML 2026