RL²-VLA:面向视觉-语言-动作模型的测试时缩放自适应RL隐成分引导
RL$^2$-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models
- National University of Singapore(新加坡国立大学)
- University of Toronto(多伦多大学)
- Singapore Technologies Engineering(新加坡科技工程公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对VLA模型分布外任务性能下降问题,提出基于VLA隐空间的自适应推理时引导框架RL²,仅在预测失败时激活成分引导,在SIMPLER等基准上分布外成功率最高提升17.3%,可迁移至真实世界。
AI中文摘要:
尽管视觉-语言-动作(VLA)模型具备令人印象深刻的视觉运动能力,但其在具有挑战性和分布外任务上的性能往往会下降。近期的测试时引导与缩放方法无需大量数据收集和重新训练即可提升性能,但动作样本往往仍集中在相似行为周围,因此继承了相关的失败模式。此外,现有方法在每个时间步都采用相同的干预策略,无论基础策略是否已有可能成功。为解决这些局限,我们引入RL²,一种利用VLA隐空间上强化学习的自适应推理时引导框架。首先,我们训练一个轻量级离线RL策略,该策略以从VLA动作专家提取的高表达隐空间为条件,在推理时将其流速度与冻结VLA的流速度进行组合。这种成分引导策略将大规模模仿学习的行为先验与离线RL在主导演示模式之外诱导的动作多样性相结合。我们进一步发现,推理时引导在成功和失败状态下遵循根本不同的缩放规律,表明动作多样性在基础VLA可能失败时最有益,但在可能成功时会不必要地扰动已准确的动作。基于这一见解,RL²仅在预测失败时激活成分引导。在SIMPLER和PolaRiS基准上,RL²在分布外设置中将成功率提升了最多17.3%,而消融实验和缩放研究证明了隐表示和RL训练的重要性。最后,真实世界实验表明,这些提升可迁移到模拟之外,确立RL²为VLA部署的实用且模块化的引导框架。
英文摘要:
Despite the impressive visuomotor capabilities enabled by Vision-Language-Action (VLA) models, their performance often degrades on challenging and out-of-domain tasks. Recent test-time steering and scaling methods improve performance without extensive data collection and retraining, but action samples often remain concentrated around similar behaviors and therefore inherit correlated failure modes. Moreover, existing methods apply the same intervention strategy at every timestep, regardless of whether the base policy is already likely to succeed. To address these limitations, we introduce $RL^2$, an adaptive inference-time steering framework that leverages Reinforcement Learning on VLA Latents. First, we train a lightweight offline RL policy conditioned on expressive latents extracted from the VLA action expert and compose its flow velocity with that of the frozen VLA during inference. This compositional steering strategy combines the behavioral priors of large-scale imitation learning with the action diversity induced by offline RL beyond dominant demonstration modes. We further discover that inference-time steering follows fundamentally different scaling laws under success and failure states, revealing that action diversity is most beneficial when the base VLA is likely to fail, but can unnecessarily perturb already-accurate actions when success is likely. Building on this insight, $RL^2$ activates compositional steering only when failure is predicted. Across the SIMPLER and PolaRiS benchmarks, $RL^2$ improves success rates by up to +17.3% in out-of-domain settings, while ablations and scaling studies demonstrate the importance of latent representations and RL training. Finally, real-world experiments demonstrate that these gains transfer beyond simulation, establishing $RL^2$ as a practical and modular steering framework for VLA deployment.