MEA:一种用于忠实模型解释的奖励驱动多智能体系统
MEA: A Reward-Driven Multi-Agent System for Faithful Model Explanations
- Aikyam Lab(Aikyam实验室)
- University of Virginia(弗吉尼亚大学)
- Vellore Institute of Technology(韦洛尔理工大学)
- Indian Institute of Technology, Patna(印度理工学院巴特那分校)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
针对现有模型解释工具使用门槛高、解释不忠实的问题,提出多智能体框架MEA,通过提议者与执行者智能体协作,并利用奖励驱动优化,在表格、文本和视觉模态上显著提升解释忠实性。
中文摘要 AI 辅助
近年来,机器学习(ML)模型在高风险领域得到了广泛应用,但对于依据其预测采取行动的专业人士而言,这些模型在很大程度上仍然是不透明的。虽然事后解释方法为理解模型行为提供了一种途径,但有效运用这些方法需要专业知识,而大多数领域专家并不具备:例如,导航高维输出、选择最佳解释以及综合来自不同工具的证据。为此,我们提出了MEA,一个多智能体框架,它完全消除了解释知识障碍:一个提议者(Proposer)智能体根据问题和模态选择并配置解释工具,而一个执行者(Actor)智能体则针对忠实性进行端到端优化,将输出转化为基于模型行为的自然语言解释,涵盖表格、文本和视觉模态。此外,我们引入了多样的问题类型,涵盖特征归因、反事实推理和虚假特征检测,每种类型都配有一个基于扰动的忠实性度量。我们发现,前沿的大语言模型系统性地产生不忠实的解释。通过针对忠实性奖励进行优化,并辅以模态自适应惩罚,MEA在六个数据集上始终优于事后解释器、智能体基线和闭源基线,其中奖励驱动的优化在未训练基线上带来了忠实性提升:表格+28%,文本+21%,视觉+34%。更广泛地说,我们的研究结果表明,AI智能体本身可以作为一种可扩展、适应性强的机器学习可解释性接口,开辟了一条通往自然语言可解释性的道路,这种可解释性超越了长期以来定义该领域的固定、单一用途工具。
英文摘要
Recent years have seen the employment of a plethora of machine learning (ML) models in high-stakes domains, but they remain largely opaque to the practitioners who act on their predictions. While post-hoc explanation methods offer a lens into this model behavior, wielding them effectively demands expertise most domain experts lack: navigating high-dimensional outputs, selecting the best explanations, and synthesizing evidence across disparate tools. To this end, we present MEA, a multi-agent framework that removes the explanation knowledge barrier entirely: a Proposer agent selects and configures explanation tools based on the question and modality, while an Actor agent is optimized end-to-end against faithfulness, transforming the outputs into natural language explanations grounded in model behavior across tabular, text, and vision modalities. Further, we introduce diverse question types spanning feature attribution, counterfactual reasoning, and spurious feature detection, each paired with a perturbation-based faithfulness metric. We find that frontier LLMs systematically produce unfaithful explanations. By optimizing against faithfulness rewards augmented with a modality-adaptive penalty, MEA consistently outperforms post hoc explainers, agentic, and closed-source baselines across six datasets, with reward-driven optimization yielding faithfulness gains of +28% (tabular), +21% (text), and +34% (vision) over the untrained backbone. More broadly, our findings suggest that AI agents themselves can serve as a scalable, adaptable interface to ML explainability, opening a path toward natural-language explainability that generalizes beyond the fixed, single-purpose tools that have long defined the field.