这会改变你的答案吗?通过反事实实验评估现实场景中大语言模型(LLM)行为的解释
Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments
浏览论文内容
中文总结 AI 辅助
本研究提出CHIVE流程,通过反事实实验评估大语言模型行为的解释,发现常见可解释性技术无预测提升,且CHIVE生成的训练数据可实现分布外泛化。
中文摘要 AI 辅助
人工智能研究的诸多领域,例如语言模型可解释性和思维链忠实性,都致力于解释模型行为,但何种解释才算“良好”呢?本研究从反事实可模拟性的角度评估解释,即该解释是否有助于预测模型在相关反事实输入上的行为。为此,我们提出了CHIVE(Counterfactual Hypothesis Investigation Via Edits,通过编辑进行反事实假设研究),这是一种新型智能体流程,可识别现实场景中模型的意外行为,并通过反事实提示编辑对其进行研究,从而为自然发生的模型行为生成数千个高质量解释及配套的反事实证据。我们以两种方式应用CHIVE:其一,评估常见的大语言模型可解释性技术是否能提升智能体预测反事实模型行为的能力,令人惊讶的是,所研究的所有可解释性技术均未产生提升;其二,我们使用CHIVE生成训练数据,发现训练模型预测CHIVE生成的反事实实验结果可泛化至各种分布外场景。总体而言,CHIVE能自动发现自然发生的大语言模型行为的解释,使我们得以评估并改进用于解释大语言模型行为的方法。
英文摘要
Many areas of AI research, such as language model interpretability and chain of thought faithfulness, seek to explain model behaviors. But what constitutes a "good" explanation? In this work, we evaluate explanations through the lens of counterfactual simulatability-whether the explanation is useful for predicting model behaviors on related counterfactual inputs. To this end, we introduce CHIVE (Counterfactual Hypothesis Investigation Via Edits), a novel agentic pipeline that identifies unexpected model behaviors in the wild and investigates them with counterfactual prompt edits. This yields thousands of high-quality explanations for naturally-occurring model behaviors along with supporting counterfactual evidence. We apply CHIVE in two ways. First, we evaluate whether common LLM interpretability techniques improve an agent's ability to predict counterfactual model behaviors. Surprisingly, we find no uplift from any of the interpretability techniques studied. Second, we use CHIVE to generate training data. We find that training models to predict outcomes of CHIVE-generated counterfactual experiments generalizes to various out-of-distribution settings. Overall, CHIVE automatically discovers explanations of naturally-occurring LLM behaviors, enabling us to evaluate and improve methods for explaining LLM behaviors.
发表机构
- Anthropic
机构由 AI 辅助整理,请以论文原文为准。