伦理决策头:基于人类反馈的强化学习在自动驾驶中实现规范伦理
The Ethical Decision Head: Operationalizing Normative Ethics in Autonomous Vehicles via Reinforcement Learning from Human Feedback
浏览论文内容
中文总结 AI 辅助
本文提出伦理决策头(EDH)框架,结合PPO与人类偏好奖励模型,在CARLA仿真中训练自动驾驶智能体,发现人类对自动驾驶伦理的理论规定与实践奖励存在差异。
中文摘要 AI 辅助
随着自动驾驶(AV)达到SAE International定义的4级和5级运行能力,其车载决策系统不仅需处理安全关键的行驶操作,还需应对随之而来的道德权重问题。本文详细介绍了伦理决策头(EDH),这是一种深度强化学习(RL)框架,将伦理推理编码为可微分的奖励信号,使策略梯度智能体在与CARLA仿真环境对齐的状态表示场景中学习符合道德规范的驾驶行为。本文实例化并评估了两种规范框架:将总伤亡最小化的功利主义框架,以及将维持行驶路线作为绝对命令的康德主义框架。EDH通过近端策略优化(PPO)训练,训练过程中使用了从200个碰撞迫近场景的成对人类偏好标注中学习到的Bradley-Terry奖励模型。结果显示,在人类监督下,规范伦理框架的可学习性存在不对称性:康德主义条件在代码本下简化为常数预测任务,起到了流程控制作用,确认了训练稳定性并排除了基础设施故障对功利主义结果的解释;而功利主义智能体学到了更令人不安的结果——人类评分者更偏好自我牺牲而非伤亡最小化,模型忠实地学习了这一偏好。人类在理论上规定的内容与实践中奖励的内容之间的这种差异表明,基于人类反馈的强化学习(RLHF)并未学习哲学家所定义的伦理,而是人类实际践行的伦理。
英文摘要
As autonomous vehicles (AVs) approach Level 4 and Level 5 operational capability [SAE International, 2018], their on- board decision systems must handle not only safety-critical locomotion but also their subsequent moral weight. This paper details the Ethical Decision Head (EDH), a deep re- inforcement learning (RL) framework that encodes ethical reasoning as a differentiable reward signal, enabling a pol- icy gradient agent to learn morally-aligned driving behavior in scenarios whose state representation is aligned with the CARLA simulation environment [Dosovitskiy et al., 2017]. Two normative frameworks are instantiated and evaluated: a Utilitarian framework minimizing total casualties and a Kan- tian framework enforcing course maintenance as a categori- cal imperative. The EDH is trained via Proximal Policy Op- timization (PPO) [Schulman et al., 2017] against a Bradley- Terry reward model [Bradley and Terry, 1952] learned from pairwise human preference annotations over 200 collision- imminent scenarios. Results reveal an asymmetry in the learnability of normative ethical frameworks under human su- pervision. The Kantian condition, which reduces to a con- stant prediction task under the codebook, serves as a pipeline control: it confirms training stability and rules out infrastruc- ture failure as an explanation for the utilitarian result. The Utilitarian agent learned something more unsettling: human raters rewarded self-sacrifice over casualty minimization, and the model learned that preference faithfully. This divergence between what humans prescribe in theory and what they re- ward in practice suggests that RLHF does not learn ethics as philosophers define it, but as humans live it.
发表机构
- Stony Brook University(石溪大学)
机构由 AI 辅助整理,请以论文原文为准。