发表机构
IIT Guwahati; NIT Agartala(印度理工学院古瓦哈提分校; 特里普拉国立技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对数据中心容器放置需同时优化功耗与亲和性并满足硬约束的问题,提出基于LLM的多智能体协作框架,通过辩论、TGA仲裁和MIR优化,在Google集群数据上实现功耗降低8.36%、亲和性提升39.30%。
AI 中文摘要
数据中心中的容器放置必须同时最小化功耗并最大化亲和性偏好,同时满足多资源容量和反亲和性约束。传统方法通常依赖固定规则,缺乏对动态集群状态的适应性,且难以扩展用于自适应决策。另一方面,元启发式方法虽然更灵活,但通常计算成本高、速度较慢,并且容易陷入局部最优。在这项工作中,我们提出了一种基于LLM驱动的多智能体协作框架的自适应且高效的方法,其中四个专门智能体在每个放置步骤中以闭环ReAct循环运行。功耗智能体和亲和性智能体就相互竞争的目标进行辩论,而放置智能体使用容差门控仲裁(TGA)解决冲突。重排智能体在后放置优化阶段通过单调改进规则(MIR)进一步细化决策。所有智能体基于环境提供的确定性、可行性过滤的候选表进行推理,确保每个提议的动作本质上满足硬约束。基于使用Google集群跟踪数据(配置为100个应用和25台机器)的实验评估,所提出的框架比功耗贪婪基线消耗的功耗低8.36%,亲和性高39.30%。消融研究证实,每个架构组件——多轮辩论机制、放置智能体中的TGA以及重排智能体中的MIR——都对整体性能有显著贡献,移除任何单个组件都会降低功耗和亲和性满意度。
英文摘要
Container placement in data centers must simultaneously minimize power consumption and maximize affinity preferences, while satisfying multi-resource capacity and anti-affinity constraints. Traditional approaches typically rely on fixed rules, which lack adaptability to dynamic cluster states and are difficult to extend for adaptive decision-making. On the other hand, meta-heuristic methods, although more flexible, are often computationally expensive, slower and prone to getting trapped in local optima. In this work, we propose an adaptive and efficient approach based on an LLM-driven multi-agent collaboration framework, where four specialized agents operate in a closed-loop ReAct cycle at each placement step. A Power Consumption Agent and an Affinity Agent debate over competing objectives, while a Placement Agent resolves conflicts using Tolerance-Gated Arbitration (TGA). A Rearrangement Agent further refines decisions through the Monotone Improvement Rule (MIR) in a post-placement refinement phase. All the agents reason over deterministic, feasibility-filtered candidate tables provided by the environment, ensuring that every proposed action inherently satisfies hard constraints. Based on experimental evaluation using the Google Cluster Trace with a configuration of 100 applications and 25 machines, the proposed framework consumes 8.36% less power and achieves 39.30% higher affinity than the power-greedy baseline. Ablation studies confirm that each architectural component, the multi-round debate mechanism, the TGA in the Placement Agent, and the MIR in the Rearrangement Agent, contributes meaningfully to the overall performance, with removal of any single component degrading both power consumption and affinity satisfaction.
CommentsAccepted in IEEE Cloud 2026
DOI:10.1109/CLOUD72782.2026.00018