ProCeedRL: Process Critic with Exploratory Demonstration Reinforcement Learning for LLM Agentic Reasoning
ProCeedRL: 带探索演示的强化学习过程批评用于LLM代理推理
机构 * Institute for Interdisciplinary Information Sciences, Tsinghua University(清华大学交叉信息研究院) ; Independent Researcher(独立研究者) ; Shanghai Qizhi Institute(上海期智研究院)
专题命中 规划推理 :reasoning(title,abstract);分类 cs.AI
AI总结 本文提出ProCeedRL,通过过程批评和探索性演示提升LLM在多轮代理任务中的推理能力,解决探索过程中错误累积问题,提高探索效率。