PhyAgentOS:一种用于具身智能体的自进化操作系统,具有解耦的认知规划和物理执行
PhyAgentOS: A Self-Evolving Operating System for Embodied Agents with Decoupled Cognitive Planning and Physical Execution
浏览论文内容
中文总结 AI 辅助
研究为具身智能体提出PhyAgentOS操作系统,其以会话为最小调度单位,用文件系统解耦认知与物理执行,通过会话验证器区分执行与任务完成,经认知记忆整合结果,有分层安全约束,在多模型和实体上验证并优于其他系统。
中文摘要 AI 辅助
视觉语言动作模型、世界模型和智能体规划器都推动了物理智能的发展,但它们的组合缺乏通用的执行抽象、共享状态、语义验证和跨异构实体的持久经验。我们提出了PhyAgentOS,一个运行时基础,将调度、验证、内存、基准测试和安全作为系统级服务提供。其以会话为中心的运行时将会话而非动作视为调度、兼容性预检、监督执行、证据收集和接受的最小单位。为了将认知与物理执行解耦,认知-物理边界是一个文件系统:状态即文件协议将跨层状态实现为带YAML的Markdown,产生可检查、可版本化的记录,且智能体和运行时层之间无代码依赖。这些视图形成一个统一的认知状态空间,使意图、能力、环境、执行和经验保持一致。会话验证器通过基于证据的成功、失败或重新规划的裁决来区分执行终止和语义任务完成。经过验证的结果通过认知记忆整合为可重用的知识和纠正经验教训,无需重新训练即可闭合试错循环。基准测试重用部署会话和验证路径,因此结果可追溯到实际执行。分层安全约束策略驱动和智能体驱动的执行:预检、动作桥接、安全防护、心跳监测和目标局部约束。验证是渐进的:游戏测试认知规划,模拟增加动力学和控制,真实机器人增加硬件噪声,认知层保持不变。PhyAgentOS在Optimus-67、StarDojo和DST-Dojo上进行了基准测试,在19个以上的模拟和物理实体上进行了验证,并在多个VLA模型上优于LIBERO、Calvin和RoboCasa365。
英文摘要
Vision-language-action models, world models, and agentic planners each advance physical intelligence, yet their composition lacks a common execution abstraction, shared state, semantic verification, and persistent experience across heterogeneous embodiments. We present PhyAgentOS, a runtime foundation delivering scheduling, verification, memory, benchmarking, and safety as system-level services. Its Session-Centered Runtime treats a session, not an action, as the minimum unit of scheduling, compatibility preflight, supervised execution, evidence collection, and acceptance. To decouple cognition from physical execution, the cognition-physics boundary is a file system: the State-as-a-File protocol materializes cross-layer state as Markdown with YAML, yielding inspectable, versionable records without code dependencies between Agent and Runtime layers. These views form a unified cognitive state space aligning intent, capabilities, environment, execution, and experience. The SessionVerifier distinguishes execution termination from semantic task completion via evidence-grounded verdicts of success, failure, or replan. Verified outcomes are consolidated through epistemic memory into reusable knowledge and corrective lessons, closing a trial-and-error loop without retraining. Benchmarking reuses the deployment session and verification path, so results trace to real execution. Layered safety constrains both policy-driven and agent-driven execution: preflight, action bridges, SafetyGuard, heartbeat monitoring, and target-local constraints. Validation is progressive: games test cognitive planning, simulation adds dynamics and control, real robots add hardware noise, with the cognitive layer held constant. PhyAgentOS is benchmarked on Optimus-67, StarDojo, and DST-Dojo, validated on 19+ simulated and physical embodiments, and gains on LIBERO, Calvin, and RoboCasa365 across multiple VLA models.
发表机构
- X-Era Lab(X-Era实验室)
- HCP Lab, Sun Yat-sen University(中山大学HCP实验室)
- Peng Cheng Laboratory(鹏城实验室)
机构由 AI 辅助整理,请以论文原文为准。