发表机构
Tata Research Development and Design Centre(塔塔研究开发设计中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出带证据门控的多智能体框架,结合图工程等技术,在 Google Cloud Platform 实现自主云 MLOps 全栈工程,可阻止无效生命周期转换,保障任务达成验证部署或可审计失败。
AI 中文摘要
各行业中,机器学习系统支撑着从预测、异常检测到预测、优化和调度等各类应用,然而将这些系统投入运行需要协调应用开发、模型流水线、云基础设施、安全、部署、监控、重训练、恢复及回滚等多项工作。我们提出一种带证据门控的多智能体框架,用于将自然语言形式的 MLOps 云工程任务转换为可验证的代码仓库及可运行的云部署。该框架结合了图工程、循环工程和智能体 harness 工程。有状态的图编排器(Graph Orchestrator)协调负责代码仓库生成、审核、执行、验证、发布和监控的专用智能体,同时管控工作流依赖、证据门、重试上限、恢复路径和终止条件。关键生命周期转换仅在其所需谓词得到可验证的执行或运行时证据支持时才会进行。验证失败会触发受限的反思、修复和重新验证,而运行时出现的失败、漂移、性能下降或策略违规证据可触发受限的适配、恢复或回滚。智能体 harness 工程通过受控能力和隔离执行环境,约束代码仓库生成、审核与修复、工件执行及云操作。我们在 Google Cloud Platform 上实现了该框架,并对代码仓库完整性、受控执行、证据门控转换、云部署上线及受限恢复进行评估。实验结果表明,该框架可阻止不被支持的生命周期转换,推动每次运行最终达成已验证的运行部署或可审计的终端失败。
英文摘要
Across industries, machine-learning systems support applications ranging from prediction and anomaly detection to forecasting, optimization, and scheduling, yet operationalizing these systems requires coordinating application development, model pipelines, cloud infrastructure, security, deployment, monitoring, retraining, recovery, and rollback. We present an evidence-gated multi-agent framework for transforming a natural-language MLOps cloud engineering task into a verified repository and operational cloud deployment. The framework combines graph engineering, loop engineering, and agent harness engineering. A stateful Graph Orchestrator coordinates specialized agents for repository generation, review, execution, verification, release, and monitoring while governing workflow dependencies, evidence gates, retry bounds, recovery paths, and termination. Consequential lifecycle transitions proceed only when their required predicates are supported by verifiable execution or runtime evidence. Verification failures activate bounded reflection, repair, and re-verification, while runtime evidence of failure, drift, degradation, or policy violation can trigger bounded adaptation, recovery, or rollback. Agent harness engineering constrains repository generation, review, and repair, artifact execution, and cloud operations through controlled capabilities and isolated execution environments. We realize the framework on Google Cloud Platform and evaluate repository completeness, controlled execution, evidence-gated transitions, cloud promotion, and bounded recovery. Our experimental results show that the framework prevents unsupported lifecycle transitions and drives each run toward either a verified operational deployment or an auditable terminal failure.
CommentsNill