视觉运动策略通过局部恢复监督的递归自我改进
Recursive Self-Improvement of Visuomotor Policies through Local Recovery Supervision
浏览论文内容
中文总结 AI 辅助
提出通过局部恢复监督递归自我改进视觉运动策略的框架,利用离线审计和教师生成纠正演示,在LIBERO-Goal和robomimic Can上分别提升成功率至88/100和112/130。
中文摘要 AI 辅助
视觉运动策略能够执行熟悉的任务,但在自身犯错后缺乏所需的纠正行为。我们提出了一个通过局部恢复监督实现递归自我改进的框架。每一轮都会审计当前策略,在受支持的失败状态生成纠正性演示,并利用这些演示来更新驱动下一轮数据收集的策略。一个离线审计器利用粗粒度和细粒度的时间证据定位未解决的失败,并指定可观察的修复目标。一个固定的多模态智能体充当使用工具的教师,通过观察、计算、执行和反馈生成恢复动作。冻结的学生模型测试每个教师端点是否支持进一步进展。如果继续失败,系统会恢复该端点并扩展演示。随后,动作级质量评估定义了具有对齐观察、质量权重和有效性掩码的连续训练窗口。仅部署学生模型。在初步的LIBERO-Goal研究中,恢复增强的后训练在100个验证场景中实现了88个成功回合,而使用相同$\pi_0$检查点进行原始数据继续训练则为78个成功回合。一项更早的基于robomimic Can的BC-RNN研究利用26个局部恢复片段将成功率从102/130提高到112/130。两次比较均匹配了2,000次额外优化步骤。
英文摘要
Visuomotor policies can execute familiar tasks yet lack the corrective behavior needed after their own mistakes. We present a framework for recursive self-improvement through local recovery supervision. Each round audits the current policy, generates corrective demonstrations at supported failure states, and uses them to update the policy that drives the next round of collection. An offline auditor locates unresolved failures using coarse and dense temporal evidence and specifies observable repair goals. A fixed multimodal agent acts as a tool-using teacher, generating recovery actions through observation, computation, execution, and feedback. The frozen student tests whether each teacher endpoint supports further progress. If continuation fails, the system restores that endpoint and extends the demonstration. Action-level quality assessment then defines continuous training windows with aligned observations, quality weights, and validity masks. Only the student is deployed. In a preliminary LIBERO-Goal study, recovery-augmented post-training achieves 88 successful episodes out of 100 validation scenes, compared with 78 for original-data continuation from the same $π_0$ checkpoint. An earlier BC-RNN study on robomimic Can improves success from 102/130 to 112/130 using 26 local recovery segments. Both comparisons match 2,000 additional optimization steps.
发表机构
- Beihang University(北京航空航天大学)
机构由 AI 辅助整理,请以论文原文为准。