arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17119cs.CL

大语言模型中反事实自我解释的实证研究

An Empirical Study of Counterfactual Self-Explanations in LLMs

  • National Technical University of Athens(雅典国立技术大学)
  • Instituto de Telecomunicações(电信研究所)

机构由 AI 辅助整理,请以论文原文为准。

Giannis Kalyvas, Giorgos Filandrianos, Orfeas Menis Mastromichalakis, Vassilis Lyberatos, Giorgos Stamou

AI总结:

本研究通过实证评估十个LLaMA-3和Qwen-2.5模型的反事实自我解释,发现模型规模是解释质量的关键,较大模型更忠实,但理由引导虽提高最小性和人类对齐,却未持续提升忠实性。

AI中文摘要:

大型语言模型可以轻松为其自身输出生成解释,但此类自我解释未必忠实于模型的行为。我们通过反事实自我解释来研究这一问题,即模型最小程度地编辑输入以使其自身预测发生变化。在情感分析和自然语言推理任务中,我们评估了来自LLaMA-3和Qwen-2.5系列的十个指令微调模型,衡量了忠实性、最小性以及与人类标注理由的一致性。我们的结果表明,模型规模是解释质量的最强决定因素:较大的模型更有可能生成能翻转自身预测并针对决策相关证据的反事实。相比之下,理由引导条件产生的编辑最小化反事实也更符合人类标注,但并未持续提高忠实性。总体而言,反事实自我解释可以提供关于模型决策的有用行为证据,但其可靠性在很大程度上取决于模型能力,应通过实证验证而非假设。

英文摘要:

Large language models can easily generate explanations for their own outputs, but such self-explanations are not necessarily faithful to the model's behavior. We study this issue through counterfactual self-explanations, where a model minimally edits an input so that its own prediction changes. Across sentiment analysis and natural language inference, we evaluate ten instruction-tuned models from the LLaMA-3 and Qwen-2.5 families, measuring faithfulness, minimality, and alignment with human-annotated rationales. Our results show that model scale is the strongest determinant of explanation quality: larger models are substantially more likely to generate counterfactuals that flip their own predictions and target decision-relevant evidence. In contrast, the rationale-guided condition produces edit-minimal counterfactuals that are also more human-aligned. However, it does not consistently improve faithfulness. Overall, counterfactual self-explanations can provide useful behavioral evidence about model decisions, but their reliability depends strongly on model capacity and should be empirically validated rather than assumed.

↑