发表机构
Southern University of Science and Technology; Beijing Zhongguancun Academy; Samsung Robotics eXperience; Wuhan University; Sun Yat-sen University(南方科技大学; 北京中关村学院; 三星机器人体验中心; 武汉大学; 中山大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
eRLT通过路由冻结VLA中的动作相关令牌与层信息构建高效状态表示,提升强化学习样本效率,在仿真与真实机器人任务中显著超越基线。
AI 中文摘要
视觉-语言-动作(VLA)模型为机器人操作提供了强大的行为先验,但如何高效地将它们适应于下游任务仍然具有挑战性。近期工作通过在线强化学习(RL)来适应冻结的VLA模型,其样本效率取决于演员和评论家所使用的状态表示的质量。现有方法要么使用与VLA无关的视觉编码器,要么通过固定压缩VLA内部表示来构建此类表示。这两种设计都没有显式地提取对下游动作细化和动作价值估计最有用的任务特定动作相关VLA特征,因此限制了样本效率。为解决这一局限,我们提出了eRLT,它通过在冻结VLA的令牌和层之间路由任务特定的动作相关信息来构建有效的状态表示。具体而言,学习到的路由令牌在多个深度动态聚合视觉-语言特征,而轻量级层路由器将这些汇总组合成固定维度的RL令牌。路由模块使用专家演示进行初始化,以捕获可预测专家动作的特征,然后使用来自在线交互的评论家反馈进行细化,用于动作价值估计。在七个LIBERO和RoboTwin任务中,与代表性基线相比,eRLT将平均归一化学习曲线AUC提升了高达23.7%。在USB连接器插入和主板带状电缆插入的真实机器人实验中,与最强基线相比,AUC分别提升了108.9%和46.7%。
英文摘要
Vision-Language-Action (VLA) models provide strong behavioral priors for robotic manipulation, yet efficiently adapting them to downstream tasks remains challenging. Recent work addresses this challenge by adapting frozen VLAs through online reinforcement learning (RL), whose sample efficiency depends on the quality of the state representation used by the actor and critic. Existing methods construct such representations either with VLA-independent visual encoders or through fixed compression of internal VLA representations. Neither design explicitly extracts the task-specific action-relevant VLA features most useful for downstream action refinement and action-value estimation, therefore limiting sample efficiency. To address this limitation, we introduce eRLT, which constructs an effective state representation by routing task-specific action-relevant information across both tokens and layers of the frozen VLA. Specifically, learned routing tokens dynamically aggregate visual-language features at multiple depths, while a lightweight layer router combines these summaries into a fixed-dimensional RL token. The routing module is initialized using expert demonstrations to capture features predictive of expert actions and then refined using critic feedback from online interactions for action-value estimation. Across seven LIBERO and RoboTwin tasks, eRLT improves mean normalized learning-curve AUC by up to 23.7% over representative baselines. Real-robot experiments on USB connector insertion and motherboard ribbon-cable insertion further show AUC improvements of 108.9% and 46.7%, respectively, over the strongest baseline.
Comments21 pages, 9 figures, 8 tables