arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AtumAI:一种用于智能体生成数据中心控制平面策略的原则性框架

AtumAI: A Principled Framework for Agentic Generation of Datacenter Control-Plane Policies

Qiushi Lin, Chaojie Zhang, Íñigo Goiri, Aditya Akella, Ricardo Bianchini, Jovan Stojkovic

arXiv 2608.02569首次发表:更新:

AI 中文总结

AtumAI是一种用于智能体生成数据中心控制平面策略的框架,通过任务编译器和进化设计循环实现形式化、可迁移、系统化的策略生成,其生成的策略在三类任务上均优于专家基线。

AI 中文摘要

数据中心的效率依赖于其控制平面策略,设计这些策略的难度日益增加:软硬件栈快速增长,设计空间庞大且相互关联,单个策略的原型设计需要数月时间。智能体AI有望实现这一搜索过程的自动化,但现成的智能体AI在三个方面存在不足:其一,它不够形式化,由于没有结构化、可搜索的问题表述,搜索过程缺乏可利用的结构,且无法保证满足硬约束;其二,它不够可迁移,每个任务都从零开始解决,因此在一个任务上的学习成果无法迁移到下一个任务;其三,它不够系统化,仅依赖大语言模型(LLM)作为唯一的候选来源,导致其仅探索设计空间的狭窄部分,容易陷入局部最优。我们提出AtumAI,这是一种采用智能体AI生成数据中心控制平面策略的框架,使该过程具备形式化、可迁移和系统化的特性。从用自然语言表述的目标出发,AtumAI会自主提出、测试并优化候选策略,直至其满足要求。该框架通过两个组件实现上述过程:一是数据中心任务编译器,它将用户请求编译为形式化、机器可检查且可搜索的任务目标、约束、决策变量及评估方法规范,实现问题形式化;二是进化设计发现循环,它在上述规范中进行搜索,通过扩散模型、进化算法和代理模型将搜索范围扩展至LLM之外。这两个组件共同将新任务的上手时间从数月的工程工作缩短至仅需编写任务描述。我们在三个具有不同问题范围、设计空间和权衡的控制平面任务上对AtumAI进行评估:工作负载放置、资源扩展和电源管理。在所有任务中,AtumAI生成的策略均始终优于专家设计的基线。

英文摘要

The efficiency of a datacenter rests on its control plane policies. Designing these policies is increasingly hard: the hardware-software stack grows fast, the design space is vast and interdependent, and prototyping a single policy takes months. Agentic AI promises to automate this search. Off the shelf, however, it falls short on three fronts. It is not formal: with no structured, searchable statement of the problem, the search has little structure to exploit and hard constraints are not guaranteed. It is not transferable: each task is solved from scratch, so nothing learned on one task carries to the next. Finally, it is not systematic: relying on the LLM as the sole source of candidates, it explores a narrow slice of the design space and settles into local optima. We introduce AtumAI, a framework that generates datacenter control-plane policies with agentic AI, making the process formal, transferable, and systematic. From a goal stated in plain language, AtumAI autonomously proposes, tests, and refines candidate policies until one satisfies the request. It does so through two components. The Datacenter Task Compiler automates problem formulation: it compiles the request into a formal, machine-checkable, and searchable specification of the task's objectives, constraints, decision variables, and evaluation methodology. The Evolutionary Design Discovery Loop then searches this specification, expanding the search beyond the LLM itself via a diffusion model, an evolutionary algorithm, and a surrogate model. Together, they reduce onboarding a new task from months of engineering to writing its description. We evaluate AtumAI on three control-plane tasks with distinct problem scopes, design spaces, and trade-offs: workload placement, resource scaling, and power management. Across all tasks, the policies generated by AtumAI consistently outperform expert-engineered baselines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑