arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

归因盲点:检索增强语言模型中源依赖的逐层轨迹诊断

The Attribution Blind Spot: Layerwise Trajectory Diagnostics for Source Reliance in Retrieval-Augmented Language Models

Zhe Yu, Wenpeng Xing, Yunzhao Wei, Bo Yang, Chen Ye, Gaolei Li, Meng Han

arXiv 2610.09493首次发表:更新:

发表机构

Zhejiang University; Binjiang Institute of Zhejiang University; East China Normal University; National FinTech Evaluation Center (Bank Card Testing Center); Hangzhou Dianzi University; Shanghai Jiaotong University(浙江大学; 浙江大学滨江研究院; 华东师范大学; 国家金融科技测评中心(银行卡检测中心); 杭州电子科技大学; 上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过受控知识冲突和潜在轨迹偏移方法,发现状态变化幅度预测源选择,而有符号PC1能选择性控制选择,揭示诊断与控制的表示分离。

AI 中文摘要

检索增强模型可以在不依赖文档的情况下匹配该文档。受控的知识冲突使源选择变得可观察,并让我们能够提出一个仅靠预测无法回答的第二个问题:哪些内部状态属性定义了有用的干预方向?我们使用潜在轨迹偏移(LTS)研究配对的隐藏状态变化,这是一种对训练拟合的第一主成分(PC1)的有符号投影,并将已验证的训练暴露与行为源选择分开。在评估的冲突中,状态变化幅度通常是更强的预测因子,而有符号的PC1是更强的选择性控制器:等范数干预改变源偏好,同时更好地保留非目标行为,且冻结方向在测试的数据集和对齐模型对之间可迁移。一项同系统的OLMo研究进一步将正向选择和控制结果与在达到的统计功效下不确定的暴露检测相结合。核心结果是一种分离:用于诊断模型将选择什么的表示,不一定是最能控制该选择的表示。

英文摘要

A retrieval-augmented model can match a document without relying on it. Controlled knowledge conflicts make source choice observable and let us ask a second question that prediction alone cannot answer: which internal-state properties define useful intervention directions? We study paired hidden-state changes with Latent Trajectory Shift (LTS), a signed projection onto a training-fitted first principal component (PC1), and keep verified training exposure separate from behavioral source choice. Across the evaluated conflicts, state-change magnitude is often the stronger predictor, whereas signed PC1 is the stronger selective controller: equal-norm interventions change source preference while better preserving non-target behavior, and the frozen direction transfers across the tested datasets and aligned model pairs. A same-system OLMo study further combines positive choice and control results with inconclusive exposure detection at the achieved power. The central result is a separation: representations that diagnose what a model will choose need not be the representations that best control that choice.

Comments27 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑