通过世界模型从人类偏好和理由中学习安全智能体行为
Learning Safe Agent Behaviour from Human Preferences and Justifications via World Models
浏览论文内容
中文总结 AI 辅助
研究在环境未知且无合适奖励函数时安全训练与部署智能体策略的问题,提出DROPJ方法,先学习世界模型,由人类在其中生成模拟轨迹并给出偏好及理由,据此训练奖励模型用于部署,实验表明该方法可降低训练成本、提升部署性能及安全性。
中文摘要 AI 辅助
我们解决在环境动态未知且无合适奖励函数的情况下安全训练智能体策略并部署良好且安全策略的问题。在安全关键环境中,我们认为传统强化学习不实用,转而借助人类输入。我们引入DROPJ,一种用于安全训练和部署的以人为本方法。首先从先前真实世界轨迹数据集学习世界模型,人类在该模型中游戏提取模拟轨迹,从中采样轨迹段对并获取人类偏好及理由,据此训练奖励模型,结合世界模型用模型预测控制直接部署智能体。通过真实用户实验发现,与其他策略相比,从用户生成信息丰富的模拟轨迹可显著降低训练计算成本,还能提高部署性能。在学习模拟器中训练时,使用偏好而非其他反馈能大幅提升部署性能。此外,偏好附带的安全理由可显著增强安全性或在部署时优先考虑与之相关的用户规定安全方面。
英文摘要
We address the problem of safely training an agent policy and deploying a good and safe policy, in settings where the environment dynamics are unknown and no suitable reward function is available. In the context of safety-critical environments, we consider traditional reinforcement learning impractical and resort to the resource of human input. We introduce DROPJ, a human-centred method for both safe training and deployment. We first learn a world model (a learned simulator) from a dataset of prior real-world trajectories. A human then plays the game in this learned simulator to extract several informative simulated trajectories. From these, we sample pairs of simulated trajectory segments and elicit from a human their preference over these segments, as well as a reason (justification) for their choice. We then train a reward model from these justified preferences and use it, together with the world model, to directly deploy the agent using model predictive control. Running real-user experiments, we find that generating informative simulated trajectories from a user significantly reduces the computational cost during training compared to other strategies, and can also improve the performance during deployment. In the context of training within a learned simulator, we show that the use of preferences rather than other types of feedback substantially improves the performance during deployment. We further demonstrate that safety justifications accompanying preferences can significantly enhance safety or prioritise user-prescribed aspects of safety associated with them during deployment.
发表机构
- University of Southampton(南安普顿大学)
- King’s College London(伦敦国王学院)
机构由 AI 辅助整理,请以论文原文为准。