发表机构
Fudan University(复旦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
IncentRL通过KL惩罚引入偏好引导并量化其对任务性能的影响,在MiniGrid上以0.01系数将成功率从90.5%提升至98%,平衡了偏好引导与原始任务目标。
AI 中文摘要
基于偏好的奖励塑形可以引导强化学习,但将偏好信号加入奖励可能会无意中改变被优化的任务。我们通过IncentRL框架解决这一问题,该框架引入偏好引导,同时明确刻画其对外部任务性能的影响。IncentRL在指定结果分布与偏好分布之间添加Kullback--Leibler(KL)惩罚。对于具有有界塑形成本的有限折扣马尔可夫决策过程,我们推导了外部价值扰动界,建立了保持原始最优策略的充分严格动作间隙条件,并通过折扣累积偏好成本刻画了大权重机制。具体示例阐明了这些保证的局限性,包括平局最优和支持不匹配。我们研究了一种实用实现,使用手工设计的基于距离的结果代理、固定偏好分布和分数加权系数搜索。在MiniGrid DoorKey-8x8上,报告的三种子平均成功率在二百万训练步后达到98%(系数为0.01),而报告的零系数基线为90.5%,同时搜索逐渐转向更小的系数。这些结果共同为基于偏好的强化学习中的核心权衡提供了原则性视角:使用额外引导来改善学习,而不过度扭曲原始任务目标。当前实验仍为描述性,尚未将KL塑形与更简单的替代方法区分开来。
英文摘要
Preference-based reward shaping can guide reinforcement learning, but adding preference signals to the reward may unintentionally change the task being optimized. We address this problem with IncentRL, a framework that introduces preference guidance while explicitly characterizing its effect on external-task performance. IncentRL adds a Kullback--Leibler (KL) penalty between a specified outcome distribution and a preferred distribution. For finite discounted Markov decision processes with bounded shaping costs, we derive an external-value perturbation bound, establish a sufficient strict-action-gap condition for preserving the original optimal policy, and characterize the large-weight regime through discounted cumulative preference cost. Exact examples clarify the limits of these guarantees, including tied optima and support mismatch. We study a practical implementation using a hand-designed, distance-based outcome proxy, a fixed preference distribution, and score-weighted coefficient search. On MiniGrid DoorKey-8x8, the reported three-seed mean success rate after two million training steps reaches 98\% with coefficient 0.01, compared with 90.5\% for the reported zero-coefficient baseline, while the search progressively shifts toward smaller coefficients. Together, these results provide a principled view of the central trade-off in preference-based RL: using additional guidance to improve learning without excessively distorting the original task objective. The current experiments remain descriptive and do not yet isolate KL shaping from simpler alternatives.