arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ExToken:用于高效视觉-语言-动作强化微调的结构化探索

ExToken: Structured Exploration for Efficient Vision-Language-Action Reinforcement Fine-tuning

Yilun Kong, Yunpeng Qing, Guozheng Ma, Haoyu Wang, Li Shen, Zhi Hou, Dacheng Tao

arXiv 2607.12931首次发表:更新:

发表机构

Nanyang Technological University; ACE Robotics; Zhejiang University(南洋理工大学; ACE机器人公司; 浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究VLA模型强化学习中探索停滞问题,提出ExToken框架,基于离线演示导出的离散行为先验调整VLA策略进行结构化探索,通过不同令牌鼓励多样行为模式,提升了探索效率与任务性能,在多任务实验中表现良好。

AI 中文摘要

强化学习(RL)在改进复杂操作任务的视觉-语言-动作(VLA)模型方面显示出巨大潜力。但其实际可扩展性因环境交互成本高而严重受限。本文首先研究了当前VLA-RL框架中的探索停滞瓶颈,发现轨迹多样性对采样效率比收集的rollout数量更重要。基于此,引入了RL探索令牌(ExToken),它基于离线演示导出的离散行为先验来调整VLA策略以进行结构化探索。在rollout收集期间通过不同令牌调整策略,鼓励智能体探索多样行为模式,提高状态-动作覆盖和探索效率。为弥合训练中的探索与部署时的确定性推理,ExToken还纳入了状态条件令牌选择器,为未见场景自适应预测有效行为模式。在模拟和现实机器人操作任务上的大量实验表明,ExToken持续加速收敛、提高任务性能并在高约束交互预算下展现出强大鲁棒性。

英文摘要

Reinforcement Learning (RL) has demonstrated significant potential for improving Vision-Language-Action (VLA) models on complex manipulation tasks. However, its practical scalability remains severely limited by the substantial cost of environmental interactions. In this work, we first investigate the exploration stagnation bottleneck in current VLA-RL frameworks and reveal that trajectory diversity is fundamentally more important to sample efficiency than the sheer quantity of collected rollouts. Motivated by these insights, we introduce RL Exploration Token (ExToken), a simple yet general framework that condition VLA policies on discrete behavioral priors derived from offline demonstrations for structured exploration. By conditioning the policy on different tokens during rollout collection, ExToken encourages the agent to explore diverse behavioral modes, substantially improving state-action coverage and exploration efficiency. To bridge exploration during training with deterministic inference at deployment, ExToken further incorporates a state-conditioned token selector that adaptively predicts effective behavioral modes for unseen scenarios. Extensive experiments across simulated and real-world robotic manipulation tasks demonstrate that ExToken consistently accelerates convergence, improves task performance, and exhibits strong robustness under highly constrained interaction budgets.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑