发表机构
School of Computer Science and Technology, Zhejiang Gongshang University; KTH Royal Institute of Technology(浙江工商大学计算机科学与技术学院; 瑞典皇家理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出TEMPO框架,通过解耦语义投影层与动作专家采用双时间尺度RL后训练,提升VLA模型在操作任务上的性能,在基准与真实任务中均优于现有方法。
AI 中文摘要
视觉-语言-动作(VLA)模型通常通过监督微调(SFT)或在线强化学习(RL)后训练适配下游操作任务。SFT易出现分布不匹配问题,现有RL方法通常对所有模型组件采用单一、统一的更新策略,忽略了它们不同的功能角色。本文提出TEMPO,一种面向VLA模型的语义-动作解耦双时间尺度RL后训练框架。TEMPO冻结预训练的视觉-语言主干网络以保留通用语义表示,并将适配范围限制在两个组件,为其配备专用RL优化循环:语义投影层和低级动作专家。我们以不同速率更新这两个组件——语义投影层更新频率低,以保持潜在动作稳定;动作专家更新频率高,以快速整合来自在线交互的控制反馈。这种解耦RL微调策略可避免快速策略更新破坏高级语义表示,同时仍能让动作专家从在线反馈中高效学习。在CALVIN基准测试和真实世界操作任务上的实验表明,TEMPO在预训练的最先进VLA模型和RL后训练基线中始终表现更优,且在两项真实世界任务上达到并维持更高的评估奖励。
英文摘要
Vision-language-action (VLA) models are commonly adapted to downstream manipulation tasks via supervised fine-tuning (SFT) or online reinforcement learning (RL) post-training. SFT is prone to distribution mismatch, and existing RL approaches typically apply a single, uniform update strategy to all model components, ignoring their distinct functional roles. We propose TEMPO, a semantic-action decoupled, two-timescale RL post-training framework for VLA models. TEMPO freezes the pretrained vision-language backbone to preserve general semantic representations, and restricts adaptation to two components with dedicated RL optimization loops: the semantic projection layer and the low-level action expert. We update them at different rates--the semantic projection layer infrequently, to keep the latent action stable, and the action expert frequently, to rapidly incorporate control feedback from online interaction. This decoupling RL fine-tuning strategy prevents fast policy updates from destabilizing high-level semantic representations while still allowing the action expert to learn efficiently from online feedback. Experiments on the CALVIN benchmark and real-world manipulation tasks demonstrate that TEMPO consistently outperforms both pretrained state-of-the-art VLA models and the RL post-training baseline, while reaching and maintaining higher evaluation rewards on two real-world tasks.