智能体操作系统(AOS):分布式智能体系统的参考操作架构
The Agent Operating System (AOS): A Reference Operating Architecture for Distributed Agentic Systems
浏览论文内容
中文总结 AI 辅助
本文提出智能体操作系统(AOS),作为分布式智能体系统的参考操作架构,用于管控异构组件以构建可管控、可靠、可观测且互操作的智能体系统。
中文摘要 AI 辅助
大型语言模型已将人工智能从孤立的预测服务转变为长期运行的分布式系统的组成部分,这些系统能够推理、调用工具、检索外部状态、委派任务并代表用户和组织采取行动。周边生态系统已推出智能体框架、工作流引擎、模型服务平台、内存系统、通信协议和可观测性工具。这些技术提升了执行能力,但未提供稳定、与实现无关的操作架构,以管控意图、选择能力、在委派时维护权限、控制不确定性、协调运行时行为以及重构重要行动的原因。本文提出智能体操作系统(Agent Operating System,AOS),这是一种针对分布式智能体系统的厂商中立参考操作架构。AOS包含两个内部层面:控制与治理层面负责意图、策略、信任、权限、置信度、可审计性、可观测性和人工监督;运行时与协调层面负责智能体生命周期、工作流协调、模型与工具路由、上下文与内存协调、调度、流量管理和运行时保障。平台服务、Linux或Windows、容器运行时以及物理基础设施仍处于AOS边界之外,通过显式接口进行集成。本文规定了AOS的概念、不变量、接口对象、优化目标、部署配置文件和可靠性责任,还指出了其中的权衡和未解决的研究问题。AOS并非现有框架或基础设施的替代品,而是作为一种操作架构,通过它可将异构组件组合成可管控、可靠、可观测且互操作的智能体系统。
英文摘要
Large language models have transformed artificial intelligence from isolated prediction services into components of long-running, distributed systems that reason, invoke tools, retrieve external state, delegate tasks, and act on behalf of users and organizations. The surrounding ecosystem has responded with agent frameworks, workflow engines, model-serving platforms, memory systems, communication protocols, and observability tools. These technologies improve execution, but they do not provide a stable, implementation-independent operating architecture for governing intent, selecting capabilities, preserving authority across delegation, controlling uncertainty, coordinating runtime behavior, and reconstructing why consequential actions occurred. This paper proposes the Agent Operating System (AOS), a vendor-neutral reference operating architecture for distributed agentic systems. AOS contains two internal planes: a Control & Governance Plane responsible for intent, policy, trust, authority, confidence, auditability, observability, and human oversight; and a Runtime & Coordination Plane responsible for agent lifecycle, workflow coordination, model and tool routing, context and memory coordination, scheduling, traffic management, and runtime assurance. Platform services, Linux or Windows, container runtimes, and physical infrastructure remain outside the AOS boundary and are integrated through explicit interfaces. The paper specifies AOS concepts, invariants, interface objects, optimization objectives, deployment profiles, and reliability responsibilities. It also identifies tradeoffs and unresolved research questions. AOS is not presented as a replacement for existing frameworks or infrastructure; it is proposed as the operating architecture through which heterogeneous components can be composed into governable, reliable, observable, and interoperable agentic systems.