从似是而非到可付诸行动:关于大语言模型自我解释的立场
From Plausible to Actionable: A Position on LLM Self-Explanations
浏览论文内容
中文总结 AI 辅助
探讨大语言模型自我解释,指出其虽看似合理但忠实性存疑。从传统可解释人工智能角度识别评估局限,提出评估合理性与忠实性的指南,并强调评估应拓展到可操作性及相关应用,以支持不同利益相关者的明智决策和适当行动。
中文摘要 AI 辅助
大语言模型(LLMs)能生成合理化自身决策的自然语言解释,即自我解释,这成为可解释人工智能的一个有前景方向。虽自我解释看似合理,但能否忠实反映模型推理过程存疑。本文指出自我解释可能似是而非却具可操作性。从传统可解释人工智能角度,识别评估LLM生成自我解释的标准协议局限,提出评估其合理性与忠实性的实用指南,还强调评估应拓展到可操作性及相关应用。
英文摘要
Large Language Models (LLMs) can generate natural language explanations that rationalize their own decisions, a phenomenon commonly referred to as self-explanations. Such explanations have emerged as a promising direction for explainable artificial intelligence (XAI), particularly for interpreting LLM behavior. However, while self-explanations often appear plausible, whether they faithfully reflect a model's underlying reasoning process remains an open question. In this opinion paper, we argue that self-explanations can be highly plausible, questionably faithful, and yet highly actionable. From a traditional XAI perspective, we identify the limitations of standard evaluation protocols for LLM-generated self-explanations and propose practical guidelines for assessing their plausibility and faithfulness. Moreover, we argue that evaluation should extend beyond these criteria to actionability, highlighting applications of LLM rationalization capabilities that support informed decision-making and appropriate action across diverse stakeholders.
发表机构
- Utrecht University(乌得勒支大学)
- Scuola Normale Superiore(高等师范学校)
- University of Pisa(比萨大学)
机构由 AI 辅助整理,请以论文原文为准。