发表机构
University of Science and Technology of China; Hefei University of Technology(中国科学技术大学; 合肥工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多智能体LLM服务的前缀缓存权衡问题,提出TOPAS调度器,在SGLang框架中实现,经实验可显著降低合成及MetaGPT工作负载下的作业完成时间。
AI 中文摘要
前缀缓存为多智能体大语言模型(LLM)服务带来了根本权衡:为智能体保留较长的系统提示键值(KV)缓存可加速后续调用,但会减少用于批量处理并发请求的GPU内存。在多阶段工作流中,现有调度器往往优先考虑即时前缀局部性或整体工作流进度。然而,在共享KV缓存预算下,单独优化任一目标都会因下游延迟或频繁前缀替换而延长任务级作业完成时间(JCT)。为平衡二者,本文提出TOPAS,即面向任务的前缀感知调度器,它共同决定保留哪些智能体前缀在缓存中、调度哪些请求执行。TOPAS通过权衡每个任务最长剩余服务路径的预期减少量与下游前缀复用的近期收益来对候选决策后状态评分,同时考虑前缀移动和抢占的成本。还引入了任务级老化机制以防止饥饿。我们在SGLang框架中实现了TOPAS,并在三个合成有向无环图(DAG)和两个MetaGPT软件开发工作流上评估其性能。与每个工作负载和指标下表现最佳的基线相比,TOPAS在合成工作负载上将平均/第99百分位JCT分别降低了多达39.8%/49.4%,在MetaGPT-SOP上降低了平均JCT 9.8%,在MetaGPT-TL上降低了平均/第99百分位JCT 22.0%/26.6%。
英文摘要
Prefix caching introduces a fundamental tradeoff in multi-agent large language model (LLM) serving: retaining a long system-prompt key-value (KV) cache for an agent accelerates future calls, yet it reduces the GPU memory available for batching concurrent requests. In multi-stage workflows, existing schedulers tend to prioritize either immediate prefix locality or overall workflow progress. However, under a shared KV cache budget, optimizing either objective in isolation can prolong tasklevel job completion time (JCT) through downstream delays or frequent prefix replacement. To strike a balance, we here propose TOPAS, a Task-Oriented Prefix-Aware Scheduler that jointly decides which agent prefixes to keep in the cache and which requests to schedule for execution. TOPAS scores candidate post-decision states by trading off the expected reduction in each task's longest remaining service path against the near-term benefit of downstream prefix reuse, accounting for the costs of prefix movement and preemption. A task-level aging mechanism is also incorporated to prevent starvation. We implement TOPAS within the SGLang framework and assess its performance on three synthetic DAGs and two MetaGPT software-development workflows. Compared with the best performing baseline for each workload and metric, TOPAS reduces the mean/p99 JCT by up to 39.8%/49.4% on the synthetic workloads, while lowering mean JCT by 9.8% on MetaGPT-SOP and mean/p99 JCT by 22.0%/26.6% on MetaGPT-TL.
Comments8 pages