通过可解释强化学习(XRL)方法对修复智能体Bug的帮助程度来评估XRL方法
Evaluating RL Explainability Methods by How Much They Help Fix Bugs in Agents
浏览论文内容
中文总结 AI 辅助
本文提出EvalXRL基准,通过大型语言模型编码智能体结合XRL方法诊断修复RL智能体故障的效果,实现对多种XRL方法的首次闭环直接对比评估。
中文摘要 AI 辅助
本篇初步论文概述了一项计划中的可解释强化学习(XRL)方法评估基准。当前评估依赖于功能基础指标,如忠实度和紧凑性,以及人类基础代理指标,如主观评分或预测准确率。我们建议通过XRL方法生成的解释在诊断和修复故障强化学习(RL)智能体方面的有效程度来评估XRL方法。我们提出EvalXRL这一基准,其中大型语言模型(LLM)编码智能体使用不同的XRL方法诊断RL智能体中预留的故障,随后修复该故障。我们提出的基准在(环境×故障×XRL方法)三元组间迭代,并利用RL智能体的奖励信号为每种XRL方法形成最终分数。编码智能体可交互使用该方法:调用XRL方法、处理其输出、形成关于故障的新假设,并调整参数再次调用该方法以测试这些假设。这种闭环结构可描述为科学方法的简化版本。部分XRL方法会提供遵循此模式的自我评估;我们提出对多种XRL方法在闭环使用场景下的首次直接对比。
英文摘要
This preliminary paper outlines a planned evaluation benchmark for Explainable Reinforcement Learning (XRL) methods. Current evaluations rely on functionally-grounded metrics like faithfulness and compactness, and on human-grounded proxies like subjective ratings or prediction accuracy. We suggest evaluating XRL methods by how effectively their generated explanations help to diagnose and fix malfunctioning reinforcement learning (RL) agents. We propose EvalXRL, a benchmark in which a Large Language Model (LLM) coding agent uses different XRL methods to diagnose a held-out malfunction in an RL agent, and then repair it. Our proposed benchmark iterates across (environment $\times$ malfunction $\times$ XRL method) tuples and uses the reward signal of the RL agents to form a final score for each XRL method. The coding agent may use the method interactively: invoke the XRL method, process its output, form new hypotheses on what is broken, and invoke the method again with parameters adjusted for testing these hypotheses. This closed-loop structure may be described as a simplified version of the scientific method. Some XRL methods provide self-evaluations that follow this pattern; we propose the first head-to-head comparison of multiple XRL methods in closed-loop usage.
发表机构
- University of California, Berkeley(加州大学伯克利分校)
- Tufts University(塔夫茨大学)
- Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。