arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Atlas:在异构集群上优化复合AI工作流的部署

Atlas: Optimizing Deployment of Compound AI Workflows on Heterogeneous Clusters

Milos Gravara, Andrija Stanisic, Stefan Nastic

arXiv 2609.04513首次发表:更新:

发表机构

TU Wien(维也纳工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Atlas是优化复合AI工作流部署的框架,通过马尔可夫准确性预测器(MAP)实现高效准确性估计,在满足SLO的同时,可降低部署成本并提升计划排序准确性。

AI 中文摘要

复合AI工作流正被越来越多地用于通过协调多个AI模型和软件组件来服务复杂的AI任务。这种方法具备部署灵活性,因为每个工作流阶段可提供不同的模型变体和资源需求,但也拓展了部署选择范围。部署工作必须选择一个执行计划,该计划为每个复合AI工作流阶段选定AI模型并将其部署到异构集群中,以满足服务水平目标(SLO)。因此,部署优化器需要估计值来比较众多候选计划并识别可行方案。系统指标通常可按每个阶段进行分析,并根据工作流拓扑结构进行组合,但准确性无法如此,因为上游阶段的误差和信息损失会影响下游阶段的准确性。现有方法要么对完整配置进行端到端分析,扩展性差;要么使用基于乘积的准确性替代模型,将各阶段视为独立个体,可能对候选计划的排序出现错误。我们提出Atlas,这是一个在SLO约束下优化复合AI部署的框架。Atlas使用MAP(马尔可夫准确性预测器),通过相邻工作流阶段之间的局部条件准确性转换来估计配置准确性。MAP将中间输出离散化为准确性桶,并根据工作流拓扑结构组合转换分析结果,无需详尽的端到端分析即可为优化器提供准确性估计。Atlas将执行计划选择表述为一个混合整数线性规划问题,在满足SLO的前提下最大化预测准确性。在四个复合AI工作流中,MAP的斯皮尔曼相关性最高可达0.947,同时相比详尽的端到端分析,将分析成本降低了多达2.6倍。在MAP的指导下,Atlas优化器选择的执行计划与神谕准确性的差距在0.03以内,同时通过异构部署将部署成本降低了多达42%。

英文摘要

Compound AI workflows are increasingly used to serve complex AI tasks by coordinating multiple AI models and software components. This approach enables deployment flexibility, as each workflow stage can expose different model variants and resource requirements, but it also expands the deployment choices. A deployment must choose an execution plan that selects AI models for each compound AI workflow stage and places them on a heterogeneous cluster in order to satisfy SLOs. Deployment optimizers therefore need estimates to compare many candidate plans and identify feasible ones. System metrics can often be profiled per stage and composed according to workflow topology, but accuracy cannot, as errors and information loss at upstream stages affect the accuracy of downstream stages. Existing approaches either profile complete configurations end to end, which scales poorly, or use product-based accuracy surrogates that treat stages as independent and can misrank candidate plans. We introduce Atlas, a framework for optimizing compound AI deployments under SLO constraints. Atlas uses MAP, a Markovian Accuracy Predictor, to estimate configuration accuracy from local conditional accuracy transitions between adjacent workflow stages. MAP discretizes intermediate outputs into accuracy buckets and composes transition profiles according to workflow topology, giving the optimizer an accuracy estimate without exhaustive end-to-end profiling. Atlas formulates execution-plan selection as a mixed-integer linear program that maximizes predicted accuracy subject to SLOs. Across four compound AI workflows, MAP achieves Spearman correlation up to 0.947 while reducing profiling cost by up to 2.6x relative to exhaustive end-to-end profiling. Guided by MAP, the Atlas optimizer selects execution plans within 0.03 of oracle accuracy while reducing deployment cost by up to 42% through heterogeneous placement.

Comments14 pages, 10 figures, 3 tables. Accepted at the IEEE/ACM Symposium on Edge Computing (SEC 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑