On the Design of KL-Regularized Policy Gradient Algorithms for LLM Reasoning
关于为LLM推理设计KL正则化策略梯度算法的设计
专题命中 复杂问题求解 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 RPG通过统一KL正则化方法和改进的梯度估计,提升LLM推理准确率,实现稳定且可扩展的强化学习算法。
Comments Published in ICLR 2026; Project Page: https://github.com/complex-reasoning/RPG