arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PhysEvo:Astra 能够行动,让它去做

PhysEvo: Astra Can Act, Let It

Wenqing Tian, Zeyu Zhang, Zhaocheng Liu, Fengwei Liu, Qiang Liu, Liang Wang

arXiv 2610.08995首次发表:更新:

发表机构

University of Chinese Academy of Sciences; Institute of Automation, Chinese Academy of Sciences; Tsinghua University(中国科学院大学; 中国科学院自动化研究所; 清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PhysEvo 提出一种围绕冻结模型的物理递归自我改进框架,通过任务与元智能体协作诊断和修订技能,无需更新权重,在仿真和真实任务中显著提升成功率。

AI 中文摘要

Astra 能够行动,但可靠的操控取决于它观察和控制世界所依赖的系统。我们提出 PhysEvo,一个围绕单一冻结模型进行物理递归自我改进(RSI)的框架。任务智能体执行机器人任务;元智能体利用由此产生的轨迹来诊断失败、修订工具和技能,并测试修正。元智能体还可以改进自身的诊断工具,因此保留的修订既支持后续行动,也支持后续的自我改进。该过程发展了关节级控制、寻求证据的观察以及可复用的操控技能,而无需更新模型权重或单独训练的行动策略。在 42 个 RoboDojo 任务中,对保留的任务特定部署版本进行留出布局评估,得到五维平均分为 68.14/100,成功率为 62.00%,相比之下,RoboDawn 的单次 Astra 智能体(我们比较中最强的已发表参考)为 47.17%。在八个对直接 Astra 具有挑战性的操控任务上,PhysEvo 达到 55.00% 的成功率,而直接 Astra 参考仅为 1.25%。将仿真演化的控制装置部署到 AgileX PiPER 上并持续进行技能修订,在五个真实世界任务的 25 次试验中,平均分为 90.60/100,成功率为 84.00%。PhysEvo 将行动的后果转化为对冻结模型如何行动和改进的持久、可测试的改变。

英文摘要

Astra can act, yet reliable manipulation depends on the system through which it observes and controls the world. We introduce PhysEvo, a framework for physical recursive self-improvement (RSI) around a single frozen model. A task agent executes robot tasks; a meta-agent uses the resulting trajectories to diagnose failures, revise tools and skills, and test corrections. The meta-agent can also improve its own diagnostic tools, so retained revisions support both later action and later self-improvement. This process develops joint-level control, evidence-seeking observation, and reusable manipulation skills without model-weight updates or a separately trained action policy. Across 42 RoboDojo tasks, held-out-layout evaluation of retained task-specific deployment versions yields a five-dimension average score of 68.14/100 and 62.00% success, compared with 47.17% for RoboDawn's one-shot Astra agent, the strongest published reference in our comparison. On eight manipulation tasks challenging direct Astra, PhysEvo achieves 55.00% success, compared with 1.25% for the direct-Astra reference. Deploying the simulation-evolved harness on AgileX PiPER and continuing skill revision yields 90.60/100 average score and 84.00% success across 25 trials on five real-world tasks. PhysEvo turns the consequences of action into persistent, testable changes to how a frozen model acts and improves.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑