Repairing Reward Functions with Feedback to Mitigate Reward Hacking
通过反馈修复奖励函数以缓解奖励黑客
机构 * Computer Science Department, Stanford University(计算机科学系, 斯坦福大学) ; School of Informatics, The University of Edinburgh(信息学院, 埃迪索恩大学)
AI总结 通过反馈修复奖励函数以缓解奖励黑客,提出PBRR方法,通过学习过渡依赖的修正项来改进代理奖励函数,从而在较少偏好下实现高性能策略。