arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29252cs.AI

用于强化微调的动态重要示例挖掘

Dynamic Important Example Mining for Reinforcement Finetuning

Haoru Tan, Sitong Wu, Yanfeng Chen, Shizhen Zhao, Yang-Tian Sun, Tianjia Liu, Chirui Chang, Shaofeng Zhang, Samm Sun, Xiuzhe Wu, Ruobing Xie, Xiaojuan Qi

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对强化微调中样本价值固定假设的缺陷,提出DIEM框架,整合梯度对齐重要性估计器与约束批量重加权方案,在多个推理基准中优于相关基线。

中文摘要 AI 辅助

强化微调(Reinforcement Fine-tuning, RFT)正被越来越多地用于增强大模型的推理能力,但其效果受限于训练数据的选择与使用方式。大多数以数据为中心的RFT方法依赖静态或启发式的样本选择,默认假设样本的价值在训练过程中固定不变,这忽略了策略学习的非平稳动态性,可能导致次优更新。我们提出动态重要示例挖掘(Dynamic Important Example Mining, DIEM),这是一个原则性强且完全自动化的框架,可使数据利用在整个RFT过程中自适应调整。DIEM在每个优化步骤中整合两个组件:(i)梯度对齐重要性估计器,能高效近似每个样本对策略改进的边际贡献;(ii)约束批量重加权方案,可在最大化聚合效用的同时,保留更新的梯度幅度以稳定优化。在多个推理基准测试中,DIEM始终优于强大的静态和动态基线。代码将通过此https链接发布。

英文摘要

Reinforcement fine-tuning (RFT) is increasingly used to strengthen the reasoning abilities of large models, yet its effectiveness is bound by how training data are selected and used. Most data-centric RFT methods rely on static or heuristic sample selection, implicitly assuming a sample's value is fixed over training. This overlooks the non-stationary dynamics of policy learning and can lead to suboptimal updates. We propose Dynamic Important Example Mining (DIEM), a principled and fully automated framework that makes data utilization adaptive throughout RFT. DIEM integrates two components into each optimization step: (i) a gradient-alignment importance estimator that efficiently approximates each sample's marginal contribution to policy improvement; and (ii) a constrained batch reweighting scheme that maximizes aggregate utility while preserving the update's gradient magnitude to stabilize optimization. Across several reasoning benchmarks, DIEM consistently outperforms strong static and dynamic baselines. The code will be released via https://github.com/hrtan/DIEM.

发表机构

  • HKU(香港大学)
  • Tencent(腾讯)
  • CUHK(香港中文大学)
  • Stanford(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑