arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14642cs.LG

基于每回合分布的伦理强化学习智能体训练与评估

Training and Evaluating Ethical Reinforcement Learning Agents on Per-Episode Distributions

Prabhjyot Singh, Majid Ghasemi, Mark Crowley

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对强化学习智能体的伦理训练与评估问题,在Craftax基准上对比四种方法,发现ESR准则下的智能体可在每回合维持违规预算,且训练评估需关注每回合分布而非均值。

中文摘要 AI 辅助

在单一奖励信号下训练的强化学习(RL)智能体会利用设计奖励与预期行为之间的差距,当我们试图赋予RL智能体伦理行为时,这一问题尤为突出。智能体可能在平均层面表现出伦理行为,却将违规行为集中在少数糟糕回合中,环境中某一回合受到伤害的实体无法通过其他回合的良好行为得到弥补。我们在开放生存基准Craftax中比较了四种训练伦理行为的方法:带终止的标量惩罚、线性多目标权重扫描、自适应拉格朗日约束,以及在预期标量回报(ESR)准则下每回合优化的非补偿效用。所有方法均在基于检测器的单一协议下评估,该协议会统计每回合中的所有违规行为且不进行审查。在平均回报对平均违规率的前沿上,四种方法表现无差异;但在每回合层面,它们的表现差异显著。在匹配平均回报时,ESR智能体在几乎所有回合中都维持了其规定的1次违规预算(最差十分位为1.04±0.07次违规),拉格朗日方法则超出了同一预算(1.14±0.03),权重扫描的最差回合违规数翻倍(2.20±0.20)。观察增强对照实验表明,这种差异源于训练目标而非智能体的观察内容,且每回合的保证在平均前沿上没有成本。当伦理违规无法在各回合间平均抵消时,我们认为训练和评估都必须针对每回合分布而非平均值。

英文摘要

Reinforcement Learning (RL) agents trained on a single reward signal exploit the gap between the designed reward and the intended behavior. This is particularly a problem when we are trying to imbue ethical behavior into RL agents. An agent can look ethical on average while concentrating its violations in a few bad episodes, and a creature in the environment harmed in one episode is not restored by good conduct in another. We compare four ways of training ethical behavior in Craftax, an open-ended survival benchmark. The four are: scalar penalties with termination, a linear multi-objective weight sweep, an adaptive Lagrangian constraint, and a non-compensatory utility optimized per episode under the Expected Scalarized Returns (ESR) criterion. All are evaluated under a single detector-based protocol that counts every violation in every episode without censoring. On the frontier of mean return against mean violation rate, the four methods are indistinguishable; per episode they separate sharply. At matched mean return, the ESR agent holds its stated budget of one violation in effectively every episode (worst-decile 1.04 +/- 0.07 violations), the Lagrangian leaks past the same budget (1.14 +/- 0.03), and the weight sweep's worst episodes double it (2.20 +/- 0.20). An observation-augmentation control attributes the separation to the training objective rather than to what the agent observes, and the per-episode guarantee costs nothing on the mean frontier. When ethical violations do not average away across episodes, we argue both training and evaluation must target the per-episode distribution rather than the mean.

补充信息

↑