Imagine-RL:基于残差置信度引导的交叉注意力用于世界模型增强的VLA强化学习
Imagine-RL: Residual-Confidence-Guided Cross-Attention for World-Model-Augmented VLA Reinforcement Learning
- Korea Advanced Institute of Science and Technology(韩国科学技术院)
- University of Science and Technology Beijing(北京科技大学)
- Shanghai Jiao Tong University(上海交通大学)
- Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
Imagine-RL通过动作条件视觉-扭矩想象增强噪声空间VLA后训练,利用残差置信度引导交叉注意力,使评论者能评估未来后果,仅用100条轨迹在真实机器人任务上平均成功率提升23.6%和60%。
AI中文摘要:
在接触丰富的操作中,可靠的动作评估需要超越当前观察,展望未来的视觉和接触后果。现有的噪声空间强化学习能高效地引导冻结的视觉-语言-动作(VLA)策略,但其评论者很大程度上忽略了这些后果。我们提出了Imagine-RL,它通过动作条件下的视觉-扭矩想象来增强噪声空间VLA后训练。对于每个候选动作块,一个冻结的视觉-扭矩潜在世界模型(VTLWM)自回归地预测紧凑的未来表示,而无需像素重建。一个当前的图像-状态-动作查询同时关注观察到的历史和预测的未来,而先前窗口的预测残差提供了逐令牌的置信度先验,以抑制不可靠的未来令牌。通过结合当前证据与预测后果,动作评论者能更好地评估候选动作并监督演员,而VLA和VTLWM保持冻结。在四个真实机器人任务中,每个任务进行50次评估试验,Imagine-RL仅使用100条RL轨迹,平均成功率比DSRL提高了23.6%,比VLA基线提高了60%。
英文摘要:
Reliable action evaluation in contact-rich manipulation requires looking beyond the current observation to future visual and contact consequences. Existing noise-space reinforcement learning efficiently steers a frozen Vision-Language-Action (VLA) policy, but its critics largely ignore these consequences. We present Imagine-RL, which augments noise-space VLA post-training with action-conditioned visual-torque imagination. For each candidate action chunk, a frozen visual-torque latent world model (VTLWM) autoregressively predicts compact future representations without pixel reconstruction. A current image-state-action query attends to observed histories and predicted futures, while previous-window prediction residuals provide token-wise confidence priors that suppress unreliable future tokens. By combining current evidence with predicted consequences, the action critic better evaluates candidate actions and supervises the actor, while the VLA and VTLWM remain frozen. Across four real-robot tasks with 50 evaluation trials per task, Imagine-RL uses only 100 RL trajectories and improves the average success rate by (23.6%) over DSRL and by (60%) over VLA baselines.