发表机构
Humboldt-Universität zu Berlin; Karlsruhe Institute of Technology(柏林洪堡大学; 卡尔斯鲁厄理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出XAI-Arena框架,利用LLM作为评判者,从多维度可扩展地评估XAI解释质量,并经人工验证其评分与人类高度相关。
AI 中文摘要
评估可解释人工智能(XAI)方法产生的解释质量仍然具有挑战性,因为现有方法往往依赖主观的人类判断,限制了研究的可重复性、可扩展性和可比性。我们考察LLM能否作为一种可重复且可扩展的机制,对XAI解释的质量进行比较评估。我们提出XAI-Arena,一个以LLM为评判者的框架,用于对XAI解释质量进行可扩展、可重复、多维且对利益相关者敏感的评估。XAI-Arena使我们能够沿多个维度比较XAI解释,即感知简单性、清晰度、任务充分性、信任校准、可操作性、透明度、忠实性和整体可解释性。随后,我们在各种数据集、机器学习模型和利益相关者角色上对XAI解释方法进行基准测试。人工验证显示,LLM生成的评分与人类评分之间存在强正相关(Spearman's rho=.693, p<.001)。总之,基于LLM的评估能够捕捉XAI解释质量中的系统性差异,并为XAI解释的比较评估提供一个可扩展且可重复的框架。
英文摘要
Evaluating the quality of explanations produced by explainable AI (XAI) methods remains challenging because existing approaches often rely on subjective human judgment, limiting reproducibility, scalability, and comparability between studies. We examine whether LLMs can serve as a reproducible and scalable mechanism to make comparative assessments of the quality of XAI explanations. We introduce XAI-Arena, an LLM-as-a-judge framework for scalable, reproducible, multidimensional, and stakeholder-sensitive evaluation of XAI explanation quality. XAI-Arena then allows us to compare XAI explanations along various dimensions, namely, perceived simplicity, clarity, task adequacy, trust calibration, actionability, transparency, faithfulness, and overall interpretability. We then benchmark XAI explanation methods across various datasets, machine learning models, and stakeholder personas. Human validation shows a strong positive association between LLM-generated and human ratings (Spearman's rho=.693, p<.001). Together, LLM-based evaluations can capture systematic differences in XAI explanation quality and provide a scalable and reproducible framework for comparative assessment of XAI explanations.