基于大语言模型反馈的策略不变奖励塑形:混合强化学习智能体框架
Policy-Invariant Reward Shaping from LLM Feedback: A Framework for Hybrid RL Agents
另 1 家 · 查看机构详情
- AI Research and Innovation Nexus for Africa (AIRINA Labs)(非洲AI研究与创新中心(AIRINA实验室))
- Sefako Makgatho Health Sciences University (SMU)(塞法科·马加托健康科学大学)
- African Center for Advanced Studies (ACAS)(非洲高级研究中心)
- African Institute for Mathematical Sciences(非洲数学科学学院)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
该研究针对LLM衍生奖励信号的理论模糊问题,提出基于LLM反馈的策略不变奖励塑形框架,将混合架构形式化为目标增强MDP,证明其最优策略集保留保证更强,并通过数值实验验证了结果。
中文摘要 AI 辅助
将大语言模型(LLM)与强化学习(RL)结合的研究日益受到关注,但LLM衍生奖励信号的理论状态常未明确。我们将LLM规划器与RL控制器的混合架构形式化为目标增强马尔可夫决策过程(Goal-Augmented MDP),证明当LLM的每状态进度评分被用作有界势函数时,即便LLM评分不准确,所得塑形项仍能保留最优策略集,该保证比通用LLM作为奖励的方法更强。我们在小型MDP上对四种势配置(包括对抗性配置,其规模为基础奖励幅度的20倍)进行数值验证,确认了该结果。
英文摘要
Combining large language models with reinforcement learning is increasingly explored, yet the theoretical status of LLM-derived reward signals is often left implicit. We formalize the hybrid LLM-planner and RL-controller architecture as a Goal-Augmented Markov Decision Process and show that when the LLM per-state progress score is used as a bounded potential function, the resulting shaping term preserves the optimal policy set even when the LLM scores are inaccurate. This guarantee is stronger than what general LLM-as-reward approaches provide. We verify the result numerically on a small MDP under four potential configurations, including an adversarial one scaled to twenty times the base reward magnitude.