arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于临界模型的具身智能自进化学习

Policy-Level Recursive Self-Improvement for Embodied AI with a Criticality World Model

Linxuan He, Yuying Tian, Lingxiang Fan, Jiaqi Pi, Yinqiao Lu, Shang Su, Mengkai Shi, Shuo Feng

arXiv 2607.28251首次发表:更新:

AI 中文总结

本文提出基于临界模型的具身智能自进化学习方法,通过状态级临界模型引导采样易失败场景,在多类具身任务中显著降低失败率,性能优于基线及同类先进模型。

AI 中文摘要

尽管策略预训练已取得快速进展,但具身智能系统在特定任务微调阶段常陷入性能瓶颈。根本原因在于微调数据的采集方式:默认流程随机收集数据,将所有样本视为同等重要,导致数据集被常规场景主导,而对改进最具价值的罕见失败案例被遗漏。本文提出一种打破瓶颈的自进化方法,核心是从策略自身执行结果中学习状态级临界模型,用于预测未来失败概率,引导重要性采样转向易失败场景。用多样化易失败场景替代冗余常规场景后,训练中通过重要性权重重采样数据,在保持无偏学习目标的同时,从根本上提升训练池的信息密度。在四足运动、多任务操作、视觉-语言-动作基准及真实机器人任务中,该方法相比训练基线将失败率降低51%-67%,相比最先进的视觉-语言-动作模型降低8%-25%。

英文摘要

Despite advances in policy pretraining, embodied AI systems can plateau during task-specific fine-tuning as uniform scenario collection encounters fewer of the remaining failures. We address this problem with a policy-level recursive self-improvement loop: execution outcomes train a criticality world model, whose risk scores guide scenario collection for specialist training. To account for the resulting shift in scenario frequencies, weighted resampling approximately corrects the collection bias. Because specialized updates can weaken nominal behavior, a risk-based gate selects between the specialist and a frozen nominal policy at inference. We extend this procedure to reinforcement learning and behavior cloning. A sampling analysis characterizes the correction, while a controlled study examines the trade-off between critical and nominal coverage. Across five embodied domains, the composed systems reduce failure rates by 59--67% for locomotion and manipulation and 8--25% for VLA benchmarks relative to their respective baselines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑