arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SEES:一种通过失败引导的VLA策略自适应演化的具身系统

SEES: A Self-Evolving Embodied System via Failure-Guided VLA Policy Adaptation

Ziwen Li, Hanlue Zhang, Zhenyang Ren, Tianyu Huang, Runqi Lin, Haoyu Wang, Zhengqing Gao, Yandong Guo, Fakhri Karray, Tongliang Liu, Chris Russell, Mingming Gong

arXiv 2609.32698首次发表:更新:

AI 中文总结

针对VLA策略在长时程任务中因薄弱原子技能而失败且难以改进的问题,提出SEES系统,通过分解任务、监控失败、仿真中构建RL任务并在线更新共享适配器,无需专家演示即可自演化提升性能,并泛化到新任务。

AI 中文摘要

近期的视觉-语言-动作(VLA)策略在多种短时程任务上展现出良好的泛化能力。然而,它们在长时程任务上仍不可靠,部分原因在于大规模训练数据偏向于演示成本较低的单阶段操作任务。一个薄弱的原子技能可能导致多个多阶段任务失败。为解决此类失败,现有方法通常需要专家识别瓶颈并提供额外演示,这使得改进成本高昂,且在部署后可能不切实际。为此,我们提出了一种自演化具身系统(SEES),它从失败中学习,无需额外专家演示即可改进VLA策略。SEES将长时程任务分解为原子任务,并将其路由到相应的家族策略。每个家族由共享一个VLA适配器的相关原子技能组成。在执行过程中,系统自动监控原子任务结果,以识别最频繁失败的原子技能作为当前瓶颈。为克服这些瓶颈,SEES通过在仿真中恢复先前遇到的状态,并利用LLM生成任务特定的成功标准,来构建定制的强化学习任务。在线强化学习更新共享的家族适配器,以促进相关原子技能之间的正向迁移,并在演化轮次中实现累积改进。大量实验表明,SEES可与不同的VLA骨干集成,逐步提升其长时程性能。我们还观察到在未见任务上的持续改进,这为超越演化设置的迁移提供了证据。

英文摘要

Recent vision-language-action (VLA) policies demonstrate promising generalization across diverse short-horizon tasks. However, they remain unreliable on long-horizon tasks, partly because the large-scale training data is biased toward single-stage manipulation tasks that are cheaper to demonstrate. A single weak atomic skill can cause failures across multiple multi-stage tasks. To address such failures, existing methods often require experts to identify the bottleneck and provide additional demonstrations, making the improvement costly and potentially impractical after deployment. To this end, we present a Self-Evolving Embodied System (SEES) that learns from failures and improves the VLA policy without additional expert demonstrations. SEES decomposes long-horizon tasks into atomic tasks and routes them to corresponding family policies. Each family consists of related atomic skills that share one VLA adapter. During execution, the system automatically monitors atomic-task outcomes to identify the most frequently failing atomic skills as the current bottlenecks. To overcome these bottlenecks, SEES constructs tailored RL tasks in simulation by restoring previously encountered states and generating task-specific success criteria with an LLM. Online RL updates the shared family adapters to promote positive transfer among related atomic skills and cumulative improvement across evolution rounds. Extensive experiments show that SEES can be integrated with different VLA backbones to progressively improve their long-horizon performance. We also observe continued improvement on unseen tasks, providing evidence of transfer beyond the evolution settings.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑