arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

找到你做不到的事:用于自改进VLA模型的智能体真实世界强化学习

Find Something You Can't Do: Agentic Real-World Reinforcement Learning for Self-Improving VLA Models

Yuan Fang, Zechu Li, Haolei Tong, Puze Liu, Georgia Chalvatzaki

arXiv 2609.32069首次发表:更新:

AI 中文总结

针对VLA模型无法从自身失败中改进的问题,提出智能体真实世界RL框架FIND,通过场景条件化任务选择和自我评估,在8个真实操作任务中将成功率从55%提升至71.9%。

AI 中文摘要

视觉-语言-动作(VLA)模型为机器人操作提供了强大的先验知识,但通常作为冻结策略部署,无法从自身的失败中改进。真实世界强化学习(RL)提供了一条持续改进的途径,然而手动环境重置和任务成功监督阻碍了自主学习的进行。我们引入了\textbf{FIND},一个智能体真实世界RL框架,在持久工作空间中闭环了场景理解、弱点感知练习、自我评估和策略改进。FIND将自主练习重新定义为一种场景条件化、性能感知的任务选择问题:不是在每次rollout后恢复预定义场景,而是利用结果场景来决定接下来练习什么。一个视觉-语言智能体从预定义库中识别可行任务,优先选择近期成功率较低的任务,并使用成对的执行前后观测来评估结果。我们使用冻结的$\pi_{0.5}$ VLA和残差离策略RL实例化了FIND。在八个真实世界操作任务中,独立人工评估的成功率从$55\\%$提高到$71.9\\%$。一个代表性运行在6小时的交互内完成了456个自主回合,需要30次场景恢复干预,在线学习期间无需人工提供奖励标签。消融实验和系统评估进一步检验了关键设计选择、智能体评估准确性和人工干预需求。我们的网站已在以下网址公开:this http URL。

英文摘要

Vision--language--action (VLA) models provide strong priors for robotic manipulation but are typically deployed as frozen policies, unable to improve from their own failures. Real-world reinforcement learning (RL) offers a path to continued improvement, yet manual environment resets and task-success supervision hinder autonomous learning. We introduce \textbf{FIND}, an agentic real-world RL framework that closes the loop between scene understanding, weakness-aware practice, self-evaluation, and policy improvement in a persistent workspace. FIND reframes autonomous practice as a scene-conditioned, performance-aware task-selection problem: instead of restoring a predefined scene after each rollout, it uses the resulting scene to determine what to practice next. A vision--language agent identifies feasible tasks from a predefined library, prioritizes those with lower recent success rates, and evaluates outcomes using paired pre- and post-execution observations. We instantiate FIND with a frozen $π_{0.5}$ VLA and residual off-policy RL. Across eight real-world manipulation tasks, the independent human-assessed success rate improves from $55\%$ to $71.9\%$. A representative run completes 456 autonomous episodes within 6 hours of interaction, requiring 30 scene-recovery interventions and no human-provided reward labels during online learning. Ablations and systematic evaluations further examine key design choices, agent evaluation accuracy, and human intervention requirements. Our website is made publicly available at: FIND.github.io.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑