arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

哪个轨迹教会了它?BehaviorTrace 与在线强化学习中训练数据归因的局限性

Which Rollout Taught It That? BehaviorTrace and the Limits of Training-Data Attribution in Online RL

Amit Nautiyal

arXiv 2610.10422首次发表:更新:

AI 中文总结

本研究通过 BehaviorTrace 框架,在 GRPO 在线强化学习微调中评估训练数据归因方法,发现多数信号受混淆因素影响,仅触发标记梯度在行为发生处保持稳定,并提供了评估检查清单。

AI 中文摘要

当强化学习教会语言模型一种新行为时,我们能否找到教会它的训练轨迹?当一种归因方法声称能做到时,我们又如何知道答案是真实的?我们在使用 GRPO 的在线强化学习微调中研究这两个问题,采用一种具有已知原因的植入行为。我们发布了 BehaviorTrace,一个开放评估框架,它结合了全梯度草图化、植入行为设置,以及对梯度幅度、流畅度、提升空间以及跨种子和生成抽取变异性的控制。在 Qwen2.5-1.5B 的三个种子上,大部分表面上的归因信号来自混淆因素。一种仅按梯度大小对训练步骤排序、没有行为目标的控制方法,达到了随机水平的 4.2 到 4.5 倍,并在三个种子中的两个上匹配或击败了最佳目标估计器。在饱和检查点上,模型流畅度预测行为标签的能力至少与我们比较的每种梯度方法相当。一旦控制了流畅度,逐轨迹的结果会因种子和生成抽取的不同而变化,因此单次运行无法解决该问题。有一个信号在所有三个种子上都成立。触发标记的梯度与在行为实际发生位置构建的目标对齐。我们将这些发现转化为评估强化学习中归因的检查清单。我们测试了现有估计器,包括 GAS(重归一化的 TracInCP)和一种 TRAK 风格的估计器,并未提出新的估计器。

英文摘要

When reinforcement learning teaches a language model a new behavior, can we find the training rollouts that taught it? And when an attribution method says it can, how do we know the answer is real? We study both questions on online RL fine-tuning with GRPO, using a planted behavior with a known cause. We release BehaviorTrace, an open evaluation harness that combines full-gradient sketching, the planted-behavior setup, and controls for gradient magnitude, fluency, headroom, and variation across seeds and generation draws. Across three seeds on Qwen2.5-1.5B, much of the apparent attribution signal comes from confounds. A control that ranks training steps by gradient size alone, with no behavior target, reaches 4.2 to 4.5 times chance and matches or beats the best targeted estimator on two of three seeds. At saturated checkpoints, model fluency predicts the behavior label at least as well as every gradient method we compared it with. Once fluency is controlled, the per-rollout results change from seed to seed and from one generation draw to the next, so a single run cannot settle the question. One signal does hold on all three seeds. The gradient of the trigger tokens aligns with a target built where the behavior actually occurs. We turn these findings into a checklist for evaluating attribution in RL. We test existing estimators, including GAS (renormalized TracInCP) and a TRAK-style estimator, and do not propose a new one.

Comments11 pages, 2 figures, 4 tables. Code and data: https://github.com/AmitoVrito/BehaviorTrace

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑