发表机构
University of Notre Dame(圣母大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出首个提示级差分隐私下的RLVR训练方法,通过聚合梯度、单次裁剪和加噪实现隐私保护,实验表明奖励信号在MATH和GSM8K上显著提升准确率,并保留非私有GRPO的大部分收益。
AI 中文摘要
带可验证奖励的强化学习(RLVR)在可能本身是机密的问题上训练语言模型,而训练后的模型可能泄露它见过哪些问题。我们研究提示级差分隐私下的RLVR:发布的权重必须关于任何单个训练问题的存在性是(ε,δ)-差分隐私的。将单个提示的一组响应作为隐私记录,我们的方法聚合它们的梯度,对提示的贡献进行一次裁剪,添加高斯噪声,并在更新之间组合隐私损失,因此隐私预算既不依赖于每个提示的响应数量,也不依赖于裁剪范数;据我们所知,这是RLVR训练的第一个差分隐私保证。我们使用LoRA在每次运行的预算ε=8下训练Qwen2.5-1.5B-Instruct,并在相同的提示和相同的预算下,与仅移除奖励信号的对照组以及两种私有监督微调方案进行比较。奖励信号在MATH上比对照组提高了2.65个点,在GSM8K上提高了3.24个点,且在每个随机种子下均如此;该改进在格式鲁棒评分器下依然存在,在MATH上为1.3个点,且不能由响应长度解释。在相同预算下,私有模型在MATH和GSM8K上比两种监督方案高出2.3到3.8个点,保留了非私有GRPO在这些任务上85%到90%的收益,而在MATH上,八倍更紧预算的噪声最多损失1.2个点。奖励效应也延伸到CommonsenseQA,这是一个探索性的非数学任务。因此,在提示级隐私下,验证器反馈仍然是一种可用的学习信号。
英文摘要
Reinforcement learning with verifiable rewards (RLVR) trains a language model on problems that may themselves be confidential, and the trained model can reveal which problems it saw. We study RLVR under prompt-level differential privacy: the released weights must be (ε,δ)-differentially private with respect to the presence of any one training problem. Taking the group of responses to one prompt as the privacy record, our method aggregates their gradients, clips the prompt's contribution once, adds Gaussian noise, and composes the privacy loss across updates, so the budget depends on neither the number of responses per prompt nor the clipping norm; to our knowledge this is the first differential privacy guarantee for RLVR training. We train Qwen2.5-1.5B-Instruct with LoRA at a per-run budget of ε=8 and compare, on the same prompts and at the same budget, a control that removes only the reward signal and two private supervised fine-tuning recipes. The reward signal improves accuracy over the control by 2.65 points on MATH and 3.24 on GSM8K, in every seed; the improvement survives a format-robust scorer, at 1.3 points on MATH, and is not explained by response length. At the same budget the private model outperforms both supervised recipes on MATH and GSM8K by 2.3 to 3.8 points, retains 85--90% of the gain of non-private GRPO on these tasks, and on MATH the noise of an eightfold tighter budget costs at most 1.2 points. The reward effect also carries to CommonsenseQA, an exploratory non-mathematical task. Verifier feedback thus remains a usable learning signal under prompt-level privacy.