arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

小语言和视觉语言模型网络智能体中GRPO的学习率门控失败:一个可控的无效结果及其机制

A Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent: A Controlled Null and Its Mechanism

Chengguang Gan, Zhixi Cai, Yunhao Liang, Hanjun Wei, Shiwen Ni, Qinghao Zhang

arXiv 2607.12640首次发表:更新:

发表机构

Monash University; University of Chinese Academy of Sciences; Shenzhen University of Advanced Technology; Pusan National University(莫纳什大学; 中国科学院大学; 深圳先进技术大学; 釜山国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究4B到8B规模小语言和视觉语言模型网络智能体中GRPO的效果,通过控制变量实验发现其在已掌握任务上无明显提升,解释了学习率导致失败的机制,表明该耦合与模型规模相关。

AI 中文摘要

具有可验证奖励的强化学习,特别是群体相对策略优化(GRPO),现在通常在监督检查点上运行,期望产生更强的智能体。我们研究它是否能为4B到8B规模的小语言和视觉语言模型网络智能体增加技能,还是主要重塑监督模型已有的行为。在一个由18次运行组成的控制网格中,改变学习率、KL权重、种子、初始化和裁剪,没有配置能可靠地提高智能体在其已基本掌握的任务上的强监督基线成功率。在文本轨道上,中等到高学习率会使其可靠地变差。在配对测试、25个评估种子、6个训练种子、配方更改、文本和标记集截图观察以及将主干扩展到8B的情况下,无效结果仍然成立;可信的危害是文本轨道上的发现,在标记集下只是名义上的。为了表明无效结果反映的是设置而非损坏的管道,我们在奖励可通过采样获得的任务上运行相同的框架、奖励和配方,在那里成功率提高了22个百分点,配对区间不包括零。因此,GRPO只有在有提升空间时才会有帮助,即采样策略已经比贪婪策略更常成功。然后我们解释了失败原因。中等学习率会使智能体退化,高学习率会使其崩溃,这两种情况形成双重解离:嫁接将退化情况定位到注意力和MLP块,而崩溃情况无法追溯到任何单个组,主导权重移动的嵌入变化在因果上是惰性的。在4B时,后期层的有效秩在两个方向上跟踪能力;在8B时两者分离。这种耦合特定于较小的模型,所以我们将其报告为与规模相关。

英文摘要

Reinforcement learning with verifiable rewards, and Group Relative Policy Optimization (GRPO) in particular, is now run routinely on a supervised checkpoint in the hope of producing a stronger agent. We ask whether it adds skill to a small language and vision-language model web agent at the 4B to 8B scale, or whether it mostly reshapes behavior the supervised model already has. Across a control grid of 18 runs that varies learning rate, KL weight, seed, initialization, and clipping, no configuration credibly improves the success rate of a strong supervised baseline on tasks the agent has largely mastered. On the text track, moderate to high learning rates make it credibly worse. The null holds under paired testing, 25 evaluation seeds, 6 training seeds, changes to the recipe, both text and Set-of-Marks screenshot observations, and scaling the backbone to 8B; the credible harm is a text-track finding and is only nominal under Set-of-Marks. To show that the null reflects the setting and not a broken pipeline, we run the identical harness, reward, and recipe on tasks whose reward is reachable by sampling, and there the success rate rises by 22 points with a paired interval that excludes zero. GRPO therefore helps only when there is headroom to climb, meaning the sampled policy already succeeds more often than the greedy one. We then explain the failure. A middle learning rate degrades the agent and a high one collapses it, and the two regimes form a double dissociation: grafting localizes the degrade regime to the attention and MLP blocks, while the collapse regime cannot be traced to any single group, and the embedding change that dominates the weight movement is causally inert. At 4B, effective rank in the late layers tracks capability in both directions; at 8B the two come apart. This coupling is specific to the smaller model, so we report it as scale-dependent.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑