arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18008cs.LGcs.AI

基于大语言模型反馈的策略不变奖励塑形:混合强化学习智能体框架

Policy-Invariant Reward Shaping from LLM Feedback: A Framework for Hybrid RL Agents

发表机构非洲AI研究与创新中心(AIRINA实验室) · 塞法科·马加托健康科学大学 · 非洲高级研究中心
另 1 家 · 查看机构详情
  • AI Research and Innovation Nexus for Africa (AIRINA Labs)(非洲AI研究与创新中心(AIRINA实验室))
  • Sefako Makgatho Health Sciences University (SMU)(塞法科·马加托健康科学大学)
  • African Center for Advanced Studies (ACAS)(非洲高级研究中心)
  • African Institute for Mathematical Sciences(非洲数学科学学院)

机构由 AI 辅助整理,请以论文原文为准。

Christophe D. Hounwanou, John Emeka Eze, Yaé U. Gaba

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对LLM衍生奖励信号的理论模糊问题,提出基于LLM反馈的策略不变奖励塑形框架,将混合架构形式化为目标增强MDP,证明其最优策略集保留保证更强,并通过数值实验验证了结果。

中文摘要 AI 辅助

将大语言模型(LLM)与强化学习(RL)结合的研究日益受到关注,但LLM衍生奖励信号的理论状态常未明确。我们将LLM规划器与RL控制器的混合架构形式化为目标增强马尔可夫决策过程(Goal-Augmented MDP),证明当LLM的每状态进度评分被用作有界势函数时,即便LLM评分不准确,所得塑形项仍能保留最优策略集,该保证比通用LLM作为奖励的方法更强。我们在小型MDP上对四种势配置(包括对抗性配置,其规模为基础奖励幅度的20倍)进行数值验证,确认了该结果。

英文摘要

Combining large language models with reinforcement learning is increasingly explored, yet the theoretical status of LLM-derived reward signals is often left implicit. We formalize the hybrid LLM-planner and RL-controller architecture as a Goal-Augmented Markov Decision Process and show that when the LLM per-state progress score is used as a bounded potential function, the resulting shaping term preserves the optimal policy set even when the LLM scores are inaccurate. This guarantee is stronger than what general LLM-as-reward approaches provide. We verify the result numerically on a small MDP under four potential configurations, including an adversarial one scaled to twenty times the base reward magnitude.

补充信息

↑