Discretizing Reward Models
离散化奖励模型
机构 * Carnegie Mellon University(卡内基梅隆大学) ; Meta Superintelligence Labs(Meta超级智能实验室)
AI总结 针对奖励模型过度敏感导致策略不佳的问题,提出一种基于蒙特卡洛dropout的无训练离散化算法,在保持判别能力的同时降低过度敏感性,减少奖励黑客行为并提升策略效果。
作者
Natural Language Processing
离散化奖励模型
机构 * Carnegie Mellon University(卡内基梅隆大学) ; Meta Superintelligence Labs(Meta超级智能实验室)
AI总结 针对奖励模型过度敏感导致策略不佳的问题,提出一种基于蒙特卡洛dropout的无训练离散化算法,在保持判别能力的同时降低过度敏感性,减少奖励黑客行为并提升策略效果。