arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06107cs.LGcs.CL

DataFlex-RL:RLVR数据策略评估平台

DataFlex-RL: An Evaluation Platform for RLVR Data Policies

Hao Liang, Mingrui Chen, Hengyi Feng, Meiyi Qiang, Wentao Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

DataFlex-RL平台在统一GRPO配方下系统评估RLVR数据策略,发现13种配置中均匀训练表现最佳,无方法带来可复现改进。

中文摘要 AI 辅助

具有可验证奖励的强化学习(RLVR)中的数据策略决定了使用哪些rollout、如何加权以及哪些领域贡献于后续训练批次。我们引入了DataFlex-RL,一个在统一GRPO配方下比较这些选择的评估平台。我们的主要实验使用Qwen2.5-7B-Base和12个数学、逻辑及科学基准,在12个匹配的随机种子上评估了13种配置。均匀GRPO将领域平衡的平均准确率相对于未训练检查点提高了7.76个百分点。八种rollout选择或重加权方法中,没有一种达到相对于均匀采样排除零的配对95%置信区间,三种自适应混合方法中也没有一种在相同精度水平上优于固定的等量混合。在Llama-3.1-8B-Base上进行的修正12种子扩展将额外方法置于与原始对照组相同的评分尺度上,但在观察到的平均性能方面未显示出一致的优胜者。我们还通过使用数学密集型六基准摘要(包括五个数学基准和GPQA-Diamond,但不含逻辑基准)对九次Qwen2.5-7B-Instruct运行进行重新评分,并将其与领域平衡的12基准摘要进行比较,量化了评估敏感性。所得排名呈负相关,相关系数为-0.33,而保留全部12个基准的摘要则基本一致。在本文研究的受控设置中,改变数据策略可显著改变训练过程,但并未产生相对于均匀训练的可复现改进。

英文摘要

Data policies for reinforcement learning with verifiable rewards (RLVR) determine which rollouts are used, how strongly they are weighted, and which domains contribute to subsequent training batches. We introduce DataFlex-RL, an evaluation platform for comparing these choices under a common GRPO recipe. Our primary experiment evaluates 13 configurations across 12 matched seeds using Qwen2.5-7B-Base and 12 mathematics, logic, and science benchmarks. Uniform GRPO improves the domain-balanced average accuracy by 7.76 percentage points over the untrained checkpoint. None of the eight rollout-selection or reweighting methods achieves a paired 95% confidence interval that excludes zero relative to uniform sampling, and none of the three adaptive mixtures outperforms a fixed equal mixture at the same level of precision. A corrected 12-seed extension on Llama-3.1-8B-Base places the additional methods on the same score scale as the original controls, but does not reveal a consistent winner in terms of observed mean performance. We also quantify evaluation sensitivity by rescoring nine Qwen2.5-7B-Instruct runs using a math-heavy six-benchmark summary, consisting of five mathematics benchmarks and GPQA-Diamond but no logic benchmark, and comparing it with the domain-balanced 12-benchmark summary. The resulting rankings are negatively correlated, with a correlation coefficient of -0.33, whereas summaries that retain all 12 benchmarks largely agree. Across the controlled settings studied here, changing the data policy measurably changes the training process but does not produce a reproducible improvement over uniform training.

发表机构

  • Peking University(北京大学)
  • UCAS(中国科学院大学)
  • Institute for Advanced Algorithms Research, Shanghai(上海高级算法研究所)
  • Zhongguancun Academy(中关村学院)

机构由 AI 辅助整理,请以论文原文为准。

↑