arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18207cs.ROcs.LG

面向实时视觉-语言-动作策略的强化学习

Reinforcement Learning for Real-Time Vision-Language-Action Policies

  • Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

Perry Dong, Kuo-Han Hung, Dorsa Sadigh, Chelsea Finn

AI总结:

针对VLA模型推理延迟导致动作过时的问题,提出Real-Time EXPO-FT框架,通过解耦动作生成与编辑实现实时强化学习微调,在Kinetix基准和动态真实世界任务中显著提升性能。

AI中文摘要:

在大型预训练视觉-语言-动作(VLA)模型之上进行强化学习微调,为高度可靠的机器人部署提供了前景。然而,由于模型规模庞大,现代VLA模型存在高推理延迟的问题,因此用于选择动作的观测在执行时往往已过时,造成分布偏移,这可能大幅降低可靠性和性能。先前的工作探索了异步策略执行以减少延迟的影响,但这些方法大多基于模仿学习,无法提供超越训练分布以实现更高可靠性的机制。我们通过实现满足动态真实世界操作实时控制要求的强化学习微调来弥补这一差距。我们的方法基于EXPO-FT,一个用于强化学习的高样本效率、可靠的VLA微调框架,并将缓慢、表达丰富的动作生成与快速、反应式的动作编辑解耦:大型预训练VLA利用其强大的行为先验提出动作块,而轻量级编辑策略则通过根据最新观测对状态变化做出响应来编辑动作,进行快速、反应式的决策。我们将其实例化为Real-Time EXPO-FT,一个用于微调实时VLA策略的强化学习框架。在Kinetix基准上,Real-Time EXPO-FT使延迟策略在10个环境中的10个里,在延迟和非延迟方法中均达到最佳性能。在四个动态真实世界任务中,即机器人物体传递、球平衡、桌上足球踢球和动态物体抓取,在线机器人数据上限为10分钟,Real-Time EXPO-FT将平均策略性能从42%提升至97%,全程无需人工干预,展示了在挑战性真实世界动态中的快速、样本高效的适应能力。网站:此https URL

英文摘要:

Reinforcement learning fine-tuning on top of large, pretrained Vision-Language-Action (VLA) models offers promise for highly reliable robot deployment. However, because of their scale, modern VLA models suffer from high inference latency, so the observation used to select an action is often stale by execution time, creating a distribution shift that can substantially degrade reliability and performance. Prior work has explored asynchronous policy execution to reduce the effect of latency, but these methods are mostly built on imitation learning and offer no mechanism for moving beyond the training distribution toward higher reliability. We close this gap by enabling RL fine-tuning that meets the real-time control requirements of dynamic real-world manipulation. Our approach builds on EXPO-FT, a framework for sample-efficient, reliable VLA fine-tuning with reinforcement learning, and decouples slow, expressive action generation from fast, reactive action edits: a large pretrained VLA proposes action chunks using its strong behavior prior, while a lightweight edit policy performs fast, reactive decision-making by editing actions in response to changes in state, conditioned on the latest observation. We instantiate this as Real-Time EXPO-FT, an RL framework for finetuning real-time VLA policies. On the Kinetix benchmark, Real-Time EXPO-FT enables a delayed policy to achieve the best performance among delayed and non-delayed methods in 10 out of 10 environments. On four dynamic real-world tasks, robot object passing, ball balancing, table soccer kicking, and dynamic object picking, with online robot data capped at 10 minutes, Real-Time EXPO-FT improves average policy performance from 42% to 97%, all without human intervention, demonstrating rapid, sample-efficient adaptation to challenging real-world dynamics. Website: https://pd-perry.github.io/real-time-expo-ft

↑