帕累托条件强化学习的命令空间反事实解释
Command-Space Counterfactual Explanations for Pareto-Conditioned Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
该研究针对帕累托条件强化学习的帕累托条件网络,提出命令空间反事实解释方法CF-ZOO,通过边界种子方向搜索生成可操作的直观解释。
中文摘要 AI 辅助
帕累托条件网络(PCNs)通过将单一策略以期望回报命令为条件,学习多种多目标强化学习行为,但命令与状态到动作的局部映射仍不透明。我们提出针对PCNs的命令空间反事实解释:给定固定状态、原始命令和替代动作,在黑盒设置下搜索仅经最小改动的期望回报命令,使同一训练策略会选择该替代动作。我们的贡献有三:第一,将PCN解释公式化为回报命令干预,采用仅回报的PCN变体,避免了以时间范围为条件带来的额外歧义;第二,将对抗机器学习方法适配到强化学习解释中;第三,引入边界种子方向搜索,改进了命令-动作空间中纯局部优化的不足,形成我们提出的方法CF-ZOO。最终得到的解释是可操作的,且能以用户自身偏好直观表达为:“如果你的权衡略微向X偏移,智能体就会选择Y。”
英文摘要
Pareto Conditioned Networks learn multiple multi-objective reinforcement learning behaviours by conditioning a single policy on a desired return command. However, the local mapping from command and state to action remains opaque. We propose command-space counterfactual explanations for PCNs: given a fixed state, original command, and foil action, we search, in a black-box setting, for a minimally changed desired-return command under which the same trained policy would choose the foil. Our contributions are threefold. First, we formulate PCN explanations as return-command interventions, using a return-only PCN variant that avoids the added ambiguity of horizon-conditioning. Second, we adapt adversarial machine learning methods to reinforcement-learning explanations. Third, we introduce a boundary-seeded directional search that improves over purely local optimization in the command-action landscape, resulting in our proposed approach CF-ZOO. The resulting explanations are actionable and intuitively expressed in the user's own preferences: "If your trade-off had shifted slightly towards X, the agent would have chosen Y."
发表机构
- Centrum Wiskunde & Informatica (CWI)(数学和计算机科学中心(CWI))
- Eindhoven University of Technology(埃因霍温理工大学)
机构由 AI 辅助整理,请以论文原文为准。