Group Distributionally Robust Optimization-Driven Reinforcement Learning for LLM Reasoning
基于组分布鲁棒优化的强化学习用于大语言模型推理
机构 * Tencent AI Lab in Bellevue WA USA(腾讯AI实验室(西雅图华盛顿州))
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 本文提出多对手组分布鲁棒优化框架,通过动态调整训练分布提升大语言模型推理性能,实现训练后精度提升10.6%和10.1%。
Comments Keywords: Large Language Models, Reasoning Models, Reinforcement Learning, Distributionally Robust Optimization, GRPO