发表机构
IRT Saint Exupéry; IRIT, Université de Toulouse; University of Groningen; Technische Universität Berlin; BIFOLD – Berlin Institute for the Foundations of Learning and Data; Saarland University; Johannes Gutenberg-University Mainz; German Research Center for Artificial Intelligence (DFKI); Centre for European Research in Trusted AI (CERTAIN); CNRS(IRT圣埃克苏佩里; 图卢兹大学IRIT实验室; 格罗宁根大学; 柏林工业大学; 柏林学习与数据基础研究所(BIFOLD); 萨尔大学; 美因茨约翰内斯古滕贝格大学; 德国人工智能研究中心(DFKI); 欧洲可信人工智能研究中心(CERTAIN); 法国国家科学研究中心(CNRS))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文复现并扩展ConSim自动可模拟性评估,发现LLM模拟器可绕过解释直接分类或利用类别匿名化泄露标签映射,并提出改进建议。
AI 中文摘要
可模拟性是一种用于评估解释的协议,它通过衡量解释在多大程度上帮助用户预测任务模型的输出来量化其有用性。由于人工评估成本高昂,自动可模拟性用LLM模拟器取代了人类被解释者,正如ConSim(Poché等人,2025)在大规模实验中所提出的那样。我们定性复现并扩展了ConSim在测试数据集、解释家族和模拟器LLM上的解释方法排名,并发现了两个局限性。首先,当类别名称具有语义时,模拟器可以直接解决分类任务而无需依赖解释,从而获得高可模拟性。其次,类别匿名化可能奖励那些泄露隐藏标签映射的解释,这一局限性我们通过一个新的“类别即概念”基线予以揭示。这些结果与捷径假设一致:在测试设置中,模拟器预测主要依赖任务先验,而解释仅产生微小变化。我们据此提出了更稳健的自动可模拟性评估建议。
英文摘要
Simulatability is an evaluation protocol for explanations that quantifies their usefulness by how well they help a user predict a task model's outputs. Since human evaluation is costly, automated simulatability replaces human explainees with LLM simulators, as proposed in ConSim (Poché et al., 2025) for large-scale experiments. We qualitatively replicate and extend ConSim's ranking of explanation methods across the tested datasets, explanation families, and simulator LLMs, and identify two limitations. First, when class names are meaningful, simulators can obtain high simulatability by solving the classification task directly, without relying on the explanations. Second, class anonymization can reward explanations for leaking the hidden label mapping, a limitation we expose with a new classes-as-concepts baseline. These results are consistent with a shortcut hypothesis: in the tested settings, simulator predictions mainly rely on task priors, while explanations produce small changes. We derive recommendations for more robust automated simulatability evaluations.
CommentsAccepted to the BlackboxNLP 2026 Reproducibility Challenge (Special Track), EMNLP 2026