arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32416cs.ROcs.AIcs.LG

RE-0:通过局部在线蒸馏实现具身代码即策略智能体的验证递归改进

RE-0: Verified Recursive Improvement of Embodied Code-as-Policy Agents through Local On-Policy Distillation

Jiawei Zhang, Xiangrong Zhang, Rui Song, Huanbin Zhou, Chengye Song, Hongzhou Wang

首次发表
浏览论文内容

中文总结 AI 辅助

RE-0提出递归验证的改进框架,通过局部在线蒸馏利用教师修正提升代码即策略具身智能体,并证明其增益下界,实验验证了有效性。

中文摘要 AI 辅助

代码即策略智能体通过生成和执行代码来完成长时程具身任务,然而,利用比学生更强但并非全局可靠的教师来持续改进这些智能体仍是一个关键挑战。现有的蒸馏方法通常将教师的完整行为视为监督目标,因此在教师失败的狀態上错误分配训练信用。我们提出RE-0,一个递归验证的策略改进框架:RE-0不假设教师在全局上优于学生,而是请求教师对学生自身的失败历史进行局部修正,并在环境中检查每个修正是否真正有益;经过验证的修正带来即时改进。在此基础上,我们提出RE-OPD,将验证过的干预转化为在线蒸馏的监督信号。只有反事实验证过的教师干预提供分布级别的监督,按其测量的局部收益加权,并且它们引发的改进被投影回独立的学生中,因此监督应用的位置和教师获得的信用量都与学生策略共同演化。我们进一步证明,学生的每轮增益以其验证干预增益为下界,直至验证和投影误差项。在多个代码即策略具身任务上的实验表明,RE-0同时提升了教师辅助执行和独立学生的表现,并泛化到新的机器人和场景。

英文摘要

Code-as-Policy agents accomplish long-horizon embodied tasks by generating and executing code, yet continually improving them with teachers that are stronger but not globally reliable remains a key challenge. Existing distillation methods typically treat the teacher's complete behavior as the supervision target and thus misassign training credit on states where the teacher fails. We propose RE-0, a recursively verified policy improvement framework: rather than assuming that the teacher globally outperforms the student, RE-0 requests local corrections from the teacher on the student's own failure histories and checks in the environment whether each correction is genuinely beneficial; verified corrections yield immediate improvement. Building on this, we propose RE-OPD, which turns verified interventions into supervision for on-policy distillation. Only counterfactually verified teacher interventions provide distribution-level supervision, weighted by their measured local benefit, and the improvement they induce is projected back into the standalone student, so both where supervision is applied and how much credit the teacher receives co-evolve with the student policy. We further prove that the student's per-round gain is lower-bounded by its verified intervention gain up to verification and projection error terms. Experiments across multiple Code-as-Policy embodied tasks show that RE-0 improves both teacher-assisted execution and the standalone student, and generalizes to novel robots and scenes.

补充信息

↑