必要还是充分?基于行为证据评估大语言模型(LLM)的解释
Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence
浏览论文内容
中文总结 AI 辅助
该研究测试LLM解释的必要性与充分性,在两个用例中对8款模型进行评估,发现引用因素与得分相关性有限,其无法可靠识别影响最强的因素,提供了LLM解释的黑箱可靠性检查框架。
中文摘要 AI 辅助
可在智能体工作流中运行的大语言模型(LLM)决策组件通常会生成与行动相关的建议或判断,并附带解释。操作人员可能会利用这些被命名的因素来监控系统、诊断错误,或决定何时对输出进行升级处理。此类应用的前提是,解释与组件可观测的决策行为一致。我们对这些被命名因素的两种解释进行了测试:必要性,即改变某一因素会改变输出;充分性,即保留该因素并移除其他可变更信息后,输出仍能保持。我们在两个合成用例中评估了这些解释:向客户推荐顾问,以及判断提示词的有害性或风险。模型会返回一个输出以及对其影响最大的前三个因素。受控黑箱干预通过测量改变某一因素会改变输出的频率,来估算该因素的必要性得分;通过测量保留该因素会保留输出的频率,来估算充分性得分。在Claude、GPT和Gemini系列的8个模型中,顾问推荐任务中,引用的排名与必要性和充分性得分的平均斯皮尔曼相关性分别为0.349和0.354;提示词监控任务中,对应数值分别为0.431和0.580。此外,在顾问推荐响应中,有57.6%的情况下,未被引用的因素在必要性标准下得分高于得分最低的被引用因素,在充分性标准下该比例为58.1%;提示词监控任务中,对应比例分别为25.8%和8.9%。被引用的前三个因素包含有用信息,但无法可靠识别出在必要性或充分性标准下影响最强的三个因素。该框架为智能体监督中使用的解释提供了黑箱可靠性检查,同时仅适用于单个LLM决策。
英文摘要
LLM decision components that can operate within agent workflows often produce action-relevant recommendations or judgements together with explanations. Operators may use the named factors to monitor a system, diagnose errors, or decide when to escalate an output. Such use assumes that the explanations agree with the component's observable decision behaviour. We test two interpretations of the named factors: necessity, meaning that changing a factor would change the output, and sufficiency, meaning that retaining it while removing other changeable information would preserve the output. We evaluate these interpretations in two synthetic use cases: recommending advisors to clients and judging prompts for harmfulness or risk. Models return an output and the top three factors that most influenced it. Controlled black-box interventions estimate a necessity score for each factor by measuring how often changing it changes the output, and a sufficiency score by measuring how often retaining it preserves the output. Across eight models from the Claude, GPT, and Gemini families, the mean Spearman correlations between the cited ranking and the necessity and sufficiency scores are 0.349 and 0.354 for advisor recommendation, and 0.431 and 0.580 for prompt monitoring. Furthermore, an uncited factor scores above the lowest-scoring cited factor in 57.6% of advisor responses under necessity and 58.1% under sufficiency; the corresponding prompt-monitoring rates are 25.8% and 8.9%. The cited top three contain useful information but do not reliably identify the three factors with the strongest measured influence under necessity or sufficiency. The framework provides a black-box reliability check for explanations used in agent oversight while remaining scoped to individual LLM decisions.
发表机构
- BNY(纽约银行)
机构由 AI 辅助整理,请以论文原文为准。