arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

不确定性门控探索噪声抑制在线强化学习微调中流匹配视觉-语言-动作策略的任务坍缩

Uncertainty-Gated Exploration Noise Suppresses Task Collapse in Online RL Fine-Tuning of a Flow-Matching Vision-Language-Action Policy

Mehmet Turan Yardımcı, Yunus Emre Çoğurcu

arXiv 2609.28838首次发表:更新:

发表机构

Karlsruhe Institute of Technology; Çukurova University(卡尔斯鲁厄理工学院; 库库罗瓦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对在线强化学习微调流匹配VLA策略时的任务坍缩问题,提出不确定性门控探索噪声控制器,在LIBERO-10上显著优于固定和学习的噪声策略,且不依赖任务标签。

AI 中文摘要

预训练的流匹配视觉-语言-动作(VLA)策略的在线强化学习微调有望使机器人在部署后持续学习,但持续更新往往会破坏其在单个任务上的能力,而总体指标看起来仍然健康。我们在LIBERO-10上以匹配的小计算预算研究这种我们称为任务坍缩的失败模式,使用由PPO结合随机(SDE)采样训练的450M参数SmolVLA策略。三种探索噪声策略在一个实时变量上有所不同:固定噪声尺度、ReinFlow风格的学习噪声网络,以及一种不确定性门控控制器,它根据任务无关的新颖性和能力信号在任务流之间重新分配探索,无需任务标签或回合边界。在汇总定义下,固定噪声在三个种子中的两个中导致任务坍缩,学习噪声在测量到迭代200的每个种子中都发生坍缩,而控制器在其三个种子中均未发生坍缩。测量的参数位移显示控制器的动作专家持续变化,而其平均应用噪声在可用日志中接近固定尺度。匹配比较支持控制器对任务保留的效果;其跨状态和跨时间的自适应贡献未被分离。较低的固定尺度减缓了衰退但并未阻止它。在此预算下,没有任何臂优于行为克隆基线。在该结果之外还测量了该机制的两个性质,而非作为其原因:按照参考配方,训练以bfloat16运行且无fp32主副本,在此条件下动作专家的96.02%元素在连续三次迭代中保持位相同,而参考学习率下的fp32主副本在单种子观测中导致两个臂均坍缩。我们发布了在四种定义下测量每任务坍缩的工具,重新评分噪声和仪器皮重。

英文摘要

Online reinforcement learning fine-tuning of pretrained flow-matching vision-language-action (VLA) policies promises robots that keep learning after deployment, but continued updates often destroy competence on individual tasks while the aggregate still looks healthy. We study this failure mode, which we call task collapse, under a matched small-compute budget on LIBERO-10 with a 450M-parameter SmolVLA policy trained by PPO with stochastic (SDE) sampling. Three exploration-noise policies differ in one live variable: a fixed noise scale, a ReinFlow-style learned noise network, and an uncertainty-gated controller that redistributes exploration across task streams from task-agnostic novelty and competence signals, without task labels or episode boundaries. Under the pooled definition, fixed noise collapses tasks in two of three seeds and learned noise in every seed measured to iteration 200, while the controller collapses none in any of its three seeds. Measured parameter displacement shows the controller's action expert keeps changing, while its mean applied noise is close to the fixed scale in the available logs. The matched comparison supports the controller's effect on task preservation; the separate contributions of its adaptation across states and over time are not disentangled. A lower fixed scale slows the decline but does not stop it. No arm improves on the behavior-cloning baseline in this budget. Two properties of that regime are measured beside this result, not offered as its cause: following the reference recipe, training runs in bfloat16 with no fp32 master copy, under which 96.02% of the action expert's elements stay bit-identical across three consecutive iterations, and an fp32 master copy at the reference learning rate collapses both arms in a single-seed observation. We release tools measuring per-task collapse under four definitions, rescoring noise and instrument tares.

Comments38 pages, 6 figures. Submitted to ICLR 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑