Coda:利用准入灵活性优化编码智能体服务
Coda: Exploiting Admission Flexibility for Coding-Agent Serving
另 1 家 · 查看机构详情
- University of Cambridge(剑桥大学)
- SJTU(上海交通大学)
- HKUST(香港科技大学)
- Meta
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
Coda通过感知就绪状态的准入层和路由层,利用准入顺序与注意力分组的灵活性,提升编码智能体服务系统的吞吐量,平均提高输出令牌吞吐量20.3%和SLO符合吞吐量70.5%。
中文摘要 AI 辅助
由大语言模型(LLM)驱动的编码智能体在模型推理与工具调用之间反复交替,形成具有可复用键值(KV)状态和异步请求恢复的长生命周期会话。然而,逻辑就绪状态并不能确保在共享服务系统中得到高效准入。通过直接的轨迹分析和轨迹驱动重放,我们识别出两种不匹配:可复用的KV状态分布在不同的存储层级,导致状态准备的计算成本不均;而具有异构上下文长度的请求可能无法高效地共同解码。我们的核心洞察是,利用准入灵活性可以在保持请求进展的同时提升服务性能。我们提出了Coda,一个编码智能体服务系统,通过一个感知就绪状态的准入层实现这一洞察,该层包含两种机制。分层老化状态准入利用准入顺序中的有界灵活性来提高准入效率并保持请求进展。兼容性感知执行准入利用每次模型迭代中注意力分组的灵活性,以减少混合上下文干扰并提高共享解码效率。对于多工作节点配置,Coda引入了一个单独的路由层,该层考虑KV状态驻留、工作节点负载以及请求到工作节点的上下文长度兼容性,以指导高效的请求放置。我们针对不同的模型和工作负载,将Coda与最先进的编码智能体服务系统及vLLM进行了评估。在单工作节点和多工作节点实验中,Coda平均将输出令牌吞吐量和符合SLO的吞吐量分别提高了20.3%和70.5%,峰值增益分别达到29.3%和140.2%。
英文摘要
Coding agents powered by large language models (LLMs) repeatedly alternate between model inference and tool calls, creating long-lived sessions with reusable key-value (KV) states and asynchronous request resumptions. Logical readiness, however, does not ensure efficient admission in a shared serving system. Through direct trace analysis and trace-driven replay, we identify two mismatches: reusable KV states reside across storage tiers and incur unequal computational costs for state preparation, while requests with heterogeneous context lengths can decode inefficiently together. Our central insight is that exploiting admission flexibility can improve serving performance while preserving request progress. We present Coda, a coding-agent serving system that realizes this insight through a readiness-informed admission layer incorporating two mechanisms. Tiered-Aging state admission exploits bounded flexibility in admission order to improve admission efficiency and preserve request progress. Compatibility-Aware execution admission exploits flexibility in attention grouping within each model iteration to reduce mixed-context interference and improve shared decoding efficiency. For multi-worker configurations, Coda introduces a separate routing layer that considers KV-state residency, worker load, and request-to-worker context-length compatibility to guide efficient request placement. We evaluate Coda across different models and workloads against state-of-the-art coding-agent serving systems and vLLM. Across the single-worker and multi-worker experiments, Coda improves output-token and SLO-compliant throughput by 20.3% and 70.5% on average, with peak gains of 29.3% and 140.2%.