arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向人形机器人的自进化AI:部署后自我改进的机制、安全性与评估

Self-Evolving AI for Humanoids: Mechanisms, Safety, and Evaluation of Post-Deployment Self-Improvement

Loc X. Nguyen, Avi Deb Raha, Huy Q. Le, Eui-Nam Huh, Dusit Niyato, Choong Seon Hong

arXiv 2609.13236首次发表:更新:

发表机构

Kyung Hee University; G-LAMP NEXUS Institute, Kyung Hee University; Nanyang Technological University(庆熙大学; 庆熙大学G-LAMP NEXUS研究所; 南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文综述了人形机器人部署后自进化的机制、安全性与评估,提出四种进化机制及世界模型验证门约束,并指出缺乏专用基准等开放挑战。

AI 中文摘要

人形机器人正成为具身人工智能的重要组成部分,这得益于运动控制中的强化学习、预测中的世界模型以及通用控制中的视觉-语言-动作模型的进步。然而,这些系统中的大多数在部署后仍保持静态。一个策略针对固定目标进行离线训练,然后被冻结,即使任务、环境和机器人本体随时间不断漂移。自进化智能体的新兴范式旨在通过允许系统从自身的部署后经验中改进来解决这一问题。由于大多数现有研究聚焦于无实体的软件智能体,本综述考察了当智能体拥有物理身体时,自进化如何发生变化。我们首先为人形机器人定义自进化,并使用一个状态元组来表示已部署的机器人,该元组包括其策略、感知、记忆、工作流和本体。该状态由进化算子在具有终身目标的慢速外环中更新。然后,我们将文献组织为自进化的四种互补机制,按自主性递增顺序呈现:自学习、自适应、自优化和自生成。由于对人形机器人的改动可能引入物理危害,我们将安全性和不确定性视为进化算子的关键设计维度,并进一步将可容许进化形式化为由人类监督包络内的世界模型验证门所执行的约束。最后,我们提出评估应跟踪机器人的进化轨迹而非固定检查点,并指出缺乏专门为自进化人形机器人设计的基准。此外,我们概述了涵盖AI算法、车载系统和治理的开放挑战。

英文摘要

Humanoid robots are becoming an important part of embodied artificial intelligence, driven by advances in reinforcement learning for locomotion, world models for prediction, and vision-language-action models for general control. However, most of these systems remain static after deployment. A policy is trained offline for a fixed objective and then frozen, even though the tasks, environments, and robot bodies keep drifting over time. An emerging paradigm of self-evolving agents aims to address this problem by allowing systems to improve from their own post-deployment experience. Since most existing studies focus on disembodied software agents, this survey examines how self-evolution changes when an agent has a physical body. We first define self-evolution for humanoids and represent a deployed robot using a state tuple that includes its policy, perception, memory, workflow, and body. This state is updated by an evolution operator in a slow outer loop with a lifelong objective. We then organize the literature into four complementary mechanisms of self-evolution, presented in increasing order of autonomy: self-learning, self-adaptation, self-optimization, and self-generation. Since changes to a humanoid can introduce physical hazards, we treat safety and uncertainty as key design dimensions of the evolution operator, and further formulate admissible evolution as a constraint enforced by a world-model verification gate within a human-oversight envelope. Finally, we present that evaluation should track the robot's evolving trajectory rather than a fixed checkpoint, and we identify the lack of a benchmark designed specifically for self-evolving humanoids. Moreover, we outline open challenges spanning AI algorithms, on-board systems, and governance.

CommentsThe paper includes 30 pages, 9 figures, 5 tables, and is considered for publication

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑