AI 中文总结
研究大语言模型上下文归因方法在上下文与训练数据重叠时的问题,引入新评估协议和基准数据集,通过实验证明现有方法在知识重叠时归因不准确,且用于源分离时无法基于贡献分数区分。
AI 中文摘要
大语言模型(LLMs)的上下文归因方法可识别对模型响应有贡献的输入上下文。近期工作在归因上下文贡献分数方面取得初步成功。然而,当上下文与训练数据重叠时,这些方法无法区分上下文内与权重内(IW)的贡献,导致分数不可靠。基于此,本文引入:1)一种基于四个新指标(基础模型上下文归因分数(BCS)、跨模型上下文归因一致性(CAC)、归因保留分数(APS)、源分离精度(SSP))的评估协议;2)一个带有真实来源标签的基准数据集(WMDP-Cyber++),以系统评估IW重叠下的归因。实验表明,当上下文中的知识也存在于权重中时,四种著名的上下文归因方法会提供不准确的归因。最后,将这些方法用于源分离(IW与上下文内学习(ICL)),发现它们无法基于贡献分数进行区分。
英文摘要
Context attribution methods for large language models (LLMs) identify which input context contributes to the model response. Recent works show the initial success in attributing the con- tributive score of the contexts. However, we observe that when the context overlaps with the training data, these methods can- not disentangle in-context from in-weight (IW) contributions, producing unreliable scores. Based on this observation, in this work, we introduce: 1) an evaluation protocol that relies on four new metrics (base-model context attribution score (BCS), cross-model context attribution consistency (CAC), attribution preservation score (APS), source separation pre- cision (SSP)) and 2) a benchmark dataset (WMDP-Cyber++) with ground-truth provenance labels to systematically assess attribution under IW overlap. In our experiments across four well-known context attribution methods, we demonstrate that they provide unfaithful attribution when the knowledge from the context also exists in the weights. Finally, we adapt these methods for source separation (IW vs. in-context learning (ICL)) and show that they cannot do the disentanglement based on the contributive score