HELIOS:面向多云分布式系统的自主资源编排策略的护栏式LLM驱动演化
HELIOS: Guardrailed LLM-Driven Evolution of Autonomous Resource Orchestration Policies for Multi-Cloud Distributed Systems
浏览论文内容
中文总结 AI 辅助
HELIOS通过护栏式LLM驱动演化生成可执行编排策略,在离线模拟中优化多云资源编排,相比基线降低45%运营成本,并确保生产环境安全。
中文摘要 AI 辅助
在多个公有云上运行延迟敏感型服务,形成了一个任何单一提供商自动扩缩器都无法看到的优化面:按需vCPU价格因提供商而异,现货折扣和中断风险因提供商和实例类型而异,出站费用惩罚状态迁移,而提供商级故障可能使单云部署瘫痪。大型语言模型(LLM)是有吸引力的编排器,因为它们可以从自然语言环境描述中综合出非平凡的决策逻辑。然而,将LLM置于每个调度决策的关键路径上是不切实际的:对于200个服务的舰队,每次决策的推理成本将超过其管理的云容量,并为毫秒级控制循环增加多秒延迟。HELIOS将LLM移出关键路径。LLM针对固定特征接口上的可执行编排策略(小型Python程序)进行演化,并针对轨迹驱动的多云模拟器进行优化。只有冠军程序在生产中运行,并包裹在护栏中,无论演化代码的提议如何,都强制执行容量可行性、SLO类放置规则和变动预算。在真实工作负载轨迹(PlanetLab、Bitbrains、Azure)与来自AWS、Azure和GCP的2026年价格和中断频率数据结合下,演化策略相比单云最佳拟合基线将惩罚性运营成本降低了45%,相比校准的多云启发式方法降低了9-19%。它以比40%更低的成本匹配了基于预言机的MILP规划器的优质SLO性能,并优于使用2.6倍更多环境交互训练的DQN元控制器。护栏至关重要:没有它们,同一策略族在10倍现货危险压力测试下会退化到97-98%的优质层停机时间,而有护栏时为0.17%。我们发布了模拟器、策略、LLM提示/响应存档和每代适应度日志,以实现完全可复现性。
英文摘要
Operating latency-sensitive services across multiple public clouds creates an optimization surface no single provider autoscaler can see: on-demand vCPU prices differ by provider, spot discounts and interruption risks vary by provider and instance type, egress fees penalize state movement, and provider-level failures can take down single-cloud deployments. Large language models (LLMs) are appealing orchestrators because they can synthesize non-trivial decision logic from a natural-language environment description. However, placing an LLM on the critical path of every scheduling decision is impractical: for a 200-service fleet, per-decision inference would cost more than the cloud capacity it manages and add multi-second latency to a millisecond-scale control loop. HELIOS moves the LLM off the critical path. The LLM evolves executable orchestration policies, small Python programs over a fixed feature interface, against a trace-driven multi-cloud simulator. Only the champion program runs in production, wrapped in guardrails that enforce capacity feasibility, SLO-class placement rules, and churn budgets regardless of evolved code's proposal. On real workload traces (PlanetLab, Bitbrains, Azure) combined with 2026 price and interruption-frequency data from AWS, Azure, and GCP, the evolved policy reduces penalized operating cost by 45% versus a single-cloud best-fit baseline and by 9-19% versus calibrated multi-cloud heuristics. It matches an oracle-informed MILP planner's premium-SLO performance at 40% lower cost and outperforms a DQN meta-controller trained with 2.6x more environment interactions. Guardrails are essential: without them, the same policy family degrades to 97-98% premium-tier downtime under a 10x spot-hazard stress test, versus 0.17% when guarded. We release the simulator, policies, LLM prompt/response archive, and per-generation fitness logs for full reproducibility.
发表机构
- New York University(纽约大学)
- Pepperdine University(佩珀代因大学)
机构由 AI 辅助整理,请以论文原文为准。