arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.14506cs.LGcs.AI

具有可验证奖励的强化学习的非空泛化界限

Non-vacuous Generalization Bounds for Reinforcement Learning with Verifiable Rewards

Yuxuan Zhu, Rohan Alur, Daniel Kang

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对具有可验证奖励的强化学习在十亿参数规模下泛化性差的问题,通过将PAC-贝叶斯压缩界限与Gumbel-max重参数化技巧结合,并提出渐进式RLVR框架,在多领域建立非空泛化界限,性能优于基础模型且接近微调模型。

中文摘要 AI 辅助

虽然具有可验证奖励的强化学习(RLVR)被广泛用于提高大语言模型(LLMs)的推理能力,但所得模型的泛化性仍了解不足。本文在十亿参数规模下为参数高效的RLVR微调建立了首个非空泛化界限。方法是将PAC-贝叶斯压缩界限应用于此设置,并通过Gumbel-max重参数化技巧解决令牌生成的固有随机性。提出渐进式RLVR框架,集成RLVR与策略蒸馏、TinyLoRA和模型量化。实验表明该框架在四个领域产生非空泛化界限,性能优于基础模型9%-51%,且在微调模型精度的6%-11%范围内。

英文摘要

While reinforcement learning with verifiable rewards (RLVR) is widely used to improve the reasoning capabilities of large language models (LLMs), the generalizability of the resulting models remains poorly understood. In this work, we establish the first non-vacuous generalization bounds for parameter-efficient RLVR fine-tuning at the billion-parameter scale. Our approach adapts PAC-Bayes compression bounds to this setting, and addresses the inherent stochasticity of token generation by applying the Gumbel-max reparameterization trick. To operationalize these bounds, we propose the Progressive RLVR framework, which integrates RLVR with on-policy distillation, TinyLoRA, and model quantization. Progressive RLVR empirically retains 84-97% performance of standard LoRA fine-tuning while producing models that are 14,796x more compressible. We show that this framework yields non-vacuous generalization bounds in four domains: mathematical problem-solving, programming, general-knowledge reasoning, and Text-to-SQL. Our bounds exceed the accuracy of the base model by 9-51% and lie within 6-11% of the accuracy of the fine-tuned models.

发表机构

  • UIUC(伊利诺伊大学厄巴纳 - 香槟分校)
  • MIT(麻省理工学院)
  • Bridgewater AIA Labs(布里奇沃特人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑