arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.04343cs.AIcs.CL

一种基于移除的测试时提升大语言模型(LLM)忠实性的方法

A Removal Based Approach to Improve LLM Faithfulness at Test-Time

Qinglan Luo, S M A Nahian, John Guttag, S. Mazdak Abulnaga, Katie Matton

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对LLM解释的不完整性提出测试时移除未提及概念的方法,在多数据集、多模型及双指标下提升了解释忠实性,且无需修改模型参数。

中文摘要 AI 辅助

大语言模型(LLM)越来越多地用于重要决策,这使得它们的解释成为审计模型行为的重要工具。遗憾的是,这些解释可能不忠实,无法反映模型决策背后的实际推理过程。我们考虑这样一种场景:LLM会针对一个问题同时给出答案和解释。我们识别出不忠实解释的两个不同维度:不完整性,即解释遗漏了影响答案的因素;以及不健全性,即解释引用了未影响模型答案的因素。现有的提升LLM忠实性的方法包括训练时方法(需要访问模型权重和大量计算资源)和测试时方法(大多聚焦于解决不健全性问题)。我们提出一种直接针对不完整性的测试时方法:我们从输入中移除模型解释中未提及的概念,并针对缩减后的输入重新查询模型。这一操作消除了未提及的影响,同时保留了提及概念的影响。在两个数据集、多个模型家族以及两个独立的忠实性指标上,与标准提示和鼓励忠实性的提示相比,我们的方法提升了解释的忠实性。我们的方法与模型无关,可在推理时应用且无需修改模型参数,为减少隐藏影响、提升LLM辅助决策的可靠性与安全性提供了灵活机制。

英文摘要

Large language models (LLMs) are increasingly used for consequential decisions, making their explanations an important tool for auditing model behavior. Unfortunately, these explanations can be unfaithful, failing to reflect the actual reasoning underlying the model's decisions. We consider a setting in which an LLM provides both an answer and an explanation in response to a question. We identify two distinct dimensions of unfaithful explanations: incompleteness, meaning that the explanation omits factors that influence the answer, and unsoundness, meaning that the explanation cites factors that did not influence the model's answer. Existing approaches to improving LLM faithfulness include training-time methods, which require access to model weights and extensive computational resources, and test-time methods that largely focus on addressing unsoundness. We introduce a test-time approach that directly targets incompleteness. We remove from the input the concepts not credited in the model's explanation and re-query the model on the reduced input. This eliminates unmentioned influences while preserving the influence of mentioned concepts. Across two datasets, multiple model families, and two independent faithfulness metrics, our approach improves explanation faithfulness compared to both standard prompting and prompting to encourage faithfulness. Our method is model-agnostic and can be applied at inference time without modifying model parameters, providing a flexible mechanism for reducing hidden influences and improving the reliability and safety of LLM-assisted decision making.

发表机构

  • MIT CSAIL(麻省理工学院计算机科学与人工智能实验室)
  • Wellesley College(韦尔斯利学院)
  • MGH, HMS(马萨诸塞总医院、哈佛医学院)

机构由 AI 辅助整理,请以论文原文为准。

↑