发表机构
University of Luxembourg; University of Bergen(卢森堡大学; 卑尔根大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对强化学习(RL)道德智能体设计中哲学文献被边缘化的问题,提炼元规范理论的相关理念审视RL架构,以明确RL智能体道德行为的判定标准,为评估相关RL方法奠定基础。
AI 中文摘要
机器伦理学与价值对齐这两个交叉学科,致力于设计与人类价值观对齐、且以符合伦理的方式行动的人工智能体。近期这些学科出现了一种趋势:使用强化学习(RL)来设计此类智能体,而曾发挥更核心作用的哲学文献则被边缘化。在此背景下,本文追求两个目标:其一,提炼元规范理论近期研究中可用于设计人工道德智能体及价值对齐智能体的理念;其二,通过这些理念审视强化学习架构。这将为判断RL智能体的行为何时可被归类为道德行为提供更清晰的标准,同时为评估和比较不同的基于RL的机器伦理学及价值对齐方法奠定基础。
英文摘要
The overlapping disciplines of machine ethics and value alignment are concerned with designing artificial agents that are aligned with human values and that act in ethically acceptable ways. A recent trend in these disciplines is the use of reinforcement learning (RL) to design such agents, sidelining the philosophical literature that used to play a more central role. Against this backdrop, this paper pursues two goals. The first is to draw out ideas from recent work in metanormative theory that can be useful for designing artificial moral and value-aligned agents. The second is to examine the RL architecture through the lens of these ideas. This will give us clearer criteria for when an RL agent's behavior can be classified as moral, as well as a basis for evaluating and comparing different RL-based approaches to machine ethics and value alignment.
CommentsThe 14th International Workshop on Engineering Multi-Agent Systems (EMAS 2026) held May 25-26, 2026 Co-located with AAMAS 2026 Paphos, Cyprus