arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VLA强化学习的低秩结构

The Low-Rank Structure of VLA Reinforcement Learning

Minjae Oh, Yoonah Park, Jongwon Lim, Yohan Jo

arXiv 2609.34599首次发表:更新:

发表机构

Seoul National University(首尔国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究揭示VLA模型强化学习产生低秩参数更新,集中于时间步模块,其位移方向预测任务成功并可引导改进策略,为高效可解释的后训练提供新见解。

AI 中文摘要

强化学习(RL)越来越多地被用于对视觉-语言-动作(VLA)模型进行后训练,然而RL如何重塑这些策略仍然鲜为人知。我们发现,在广泛使用的基于流的VLA模型(包括π₀.₅和GR00T N1.5/N1.6)上,在LIBERO、ManiSkill、MetaWorld和CALVIN基准上进行RL会引发显著更低秩的参数更新,这些更新高度集中在动作专家的时间步模块(Timestep Modules)中,这是一个小型且此前被忽视的组件。通过系统的模块替换实验,我们进一步表明这些模块捕获了RL带来的性能提升中不成比例的部分。接着,我们刻画了这些时间步模块中编码的内容。首先,我们表明RL将它们专门化到推理过程中使用的离散去噪时间步,并且这种离散时间步训练是低秩更新的基础。其次,我们发现,在其输出中,位移向量在RL下变化最为显著,并且通过探测,我们表明位移更新方向强烈预测任务成功(ROC-AUC高达99.6%)。第三,我们发现位移更新的几何结构反映了任务关系,因为它们的成对相似性与跨任务迁移模式相关。基于这些发现,我们表明沿着位移更新方向进行引导可以进一步改进RL训练的策略,而无需额外的RL训练。总体而言,我们通过研究学习信号如何在参数空间中被编码,系统地理解了RL如何重塑VLA策略,为更高效和可解释的VLA后训练提供了见解。

英文摘要

Reinforcement learning (RL) is increasingly used to post-train vision-language-action (VLA) models, yet how RL reshapes these policies remains poorly understood. We find that RL across widely used flow-based VLA models, including $π_{0.5}$ and GR00T~N1.5/N1.6, on LIBERO, ManiSkill, MetaWorld, and CALVIN induces substantially lower-rank parameter updates that are highly concentrated in the action expert's Timestep Modules, a small and previously overlooked component. Through systematic module-replacement experiments, we further show that these modules capture a disproportionate share of the performance gains from RL. We then characterize what is encoded in these Timestep Modules. First, we show that RL specializes them to the discrete denoising timesteps used during rollouts, and that this discrete-timestep training underlies the low-rank updates. Second, we find that among their outputs, the shift vector changes most distinctly under RL, and through probing, we show that shift update directions strongly predict task success (ROC-AUC up to $99.6\%$). Third, we find that the geometry of shift updates reflects task relationships, as their pairwise similarity correlates with cross-task transfer patterns. Building on these findings, we show that steering along shift update directions further improves RL-trained policies without additional RL training. Overall, we provide a systematic understanding of how RL reshapes VLA policies by studying how learned signals are encoded in parameter space, offering insights into more efficient and interpretable VLA post-training.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑