用于智能应用的工作流感知服务层
Serving Agentic Workflows with a Physical-Plan Compiler and Adaptive Runtime
浏览论文内容
中文总结 AI 辅助
研究智能AI应用工作流,其介于现有两层之间。提出Dyserve工作流感知服务层填补空白,通过整数线性规划在异构后端池编译节点模型和验证器选择,还可在不同压力水平预求解并动态调整。
中文摘要 AI 辅助
智能AI应用形成新兴服务工作负载,请求创建工作流。现有模型服务引擎和代理框架各有局限。我们提出Dyserve,通过整数线性规划在异构后端池编译工作流节点模型和验证器选择,结合硬件吞吐量扫描,在不同压力水平预求解并在负载下动态调整,工具调用失败触发残差重新求解。
英文摘要
Efficient serving of agentic workflows requires selecting each LLM node's model, verification policy, and backend to balance output quality, latency, and throughput. These assignments must also adapt to changes in serving load. Existing approaches address parts of this problem through model routing, verifier placement, and backend scheduling. However, independent optimization overlooks their dependencies: model and backend choices determine verification cost, while verification changes the quality-cost trade-off among models. Ignoring these interactions can waste serving resources and degrade workflow performance. To address this problem, we propose \textbf{Dyserve}, which provides the missing workflow physical-planning layer between orchestration and model serving through compiler-runtime co-design. Our design is guided by three observations: planning headroom is request-dependent; node vulnerability, the impact of local errors on final correctness, depends on position and task type; and serving load changes the cost of a plan during execution. For each request, its profile-guided compiler jointly selects node implementations and prepares pressure-specialized variants for the materialized workflow. The runtime selects variants using live backend pressure and updates only undispatched assignments, without invoking the optimizer on the load-change path. Across four agentic workloads, Dyserve improves accuracy by \textbf{3-9} percentage points with \textbf{1.1-6.8}$\times$ mean-latency speedups over the highest-accuracy evaluated baseline for each workload. On a burst trace, variant switching raises the fraction of correct, on-time completions from \textbf{18.1\%} to \textbf{67.2\%} relative to admission-only execution.
发表机构
- Georgia Institute of Technology(佐治亚理工学院)
- Harvard University(哈佛大学)
- Intel(英特尔)
- Columbia University(哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。