Selective Expert Guidance for Effective and Diverse Exploration in Reinforcement Learning of LLMs
选择性专家引导实现强化学习中有效且多样的探索
机构 * School of Data Science, Fudan University(复旦大学数据科学学院) ; Shanghai Institute of Artificial Intelligence for Education, East China Normal University(华东师范大学上海智能教育研究院) ; College of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机科学与技术学院) ; Ant Group(蚂蚁集团)
专题命中 规划推理 :reasoning(abstract);分类 cs.CL、cs.AI
AI总结 提出MENTOR框架,仅在关键决策点提供专家引导,平衡探索的有效性与多样性,提升LLM在可验证奖励强化学习中的推理能力。
Comments Accepted by ICLR 2026