arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型中的证据整合

Evidence Integration in Large Language Models

Sebastien Kawada, Manolis Kellis

arXiv 2609.04290首次发表:更新:

发表机构

Massachusetts Institute of Technology; Computer Science and Artificial Intelligence Laboratory(麻省理工学院; 计算机科学与人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出分布理论解释LLMs如何整合外部证据,经千万次试验、十二种LLMs及八领域验证,发现其整合是与接收者相关的后期控制策略,且验证与整合状态可分离。

AI 中文摘要

尽管人们越来越依赖通过工具、检索增强生成、其他智能体和用户提供的外部证据进行推理的大语言模型(LLMs),但LLMs如何将此类证据整合到其已开始形成的决策中,在很大程度上仍不清楚。我们提出了一种分布理论,其中证据会根据接收者先验权重和候选证据倾向改变接收者初始答案的分布,进而产生三个预测:第一,对接收者而言更可能的候选更具说服力;第二,接收者更容易整合自身的特征错误,而非来自不同来源的外部错误;第三,相同的证据可以提升较弱的模型,却会损害较强的模型。我们通过超过一千万次试验、四个系列的十二个LLMs以及八个领域(其中四个是物理和生命科学领域的科学发现任务:量子力学、物理学、遗传学和分子生物学)验证了这些预测。该理论还产生了一个与接收者相关的可靠性前沿:接收者一致的错误比相同比例的随机错误对性能的抑制作用更显著。LLMs甚至会在内部验证候选答案无效后仍对其进行整合(在命题约束下为93%-100%;在未见过的物理和生命科学推理任务上高达99.4%),这表明证据整合是一种针对现有分布的、与接收者相关的控制策略,由接收者属性而非对证据源的标量信任决定。因果干预显示,候选整合在网络后期执行,表现为一系列结构化步骤:接纳外部候选答案、对其进行提升并将其转移到答案状态。验证的表示是可解码的,但对答案几乎没有因果影响。J-lens分解表明,口头验证背后的状态与候选整合背后的状态完全可分离。

英文摘要

Despite increasing reliance on LLMs that reason with external evidence supplied by tools, retrieval-augmented generation, other agents, and users, how LLMs integrate such evidence into decisions they have already begun to form remains largely unclear. We present a distributional theory in which evidence shifts the receiver's distribution of initial answers, driven by a receiver prior weight and a candidate evidence tilt, leading to three predictions. First, candidates more probable to the receiver are more persuasive. Second, receivers more readily integrate characteristic errors of their own than foreign errors from different sources. Third, identical evidence can improve weaker models and harm stronger ones. We confirm these over ten million trials, twelve LLMs from four families, and eight domains, four of them scientific discovery tasks in the physical and life sciences: quantum mechanics, physics, genetics, and molecular biology. The law also yields a receiver-relative reliability frontier: receiver-congruent errors depress performance more steeply than random errors of the same rate. LLMs also integrate candidates even after internally verifying their invalidity (93-100% with propositional constraints; up to 99.4% on held-out physical and life-sciences reasoning), demonstrating evidence integration is a receiver-specific control policy over existing distributions, determined by receiver properties rather than scalar trust in the evidence source. Causal interventions show candidate integration is implemented late in the network, as a structured sequence of steps admitting external candidate answers, promoting them, and transporting them into the answer state. Representations of verification are decodable but have little causal impact on answers. A J-lens decomposition shows the state underlying verbalized verification is fully dissociable from that underlying candidate integration.

Comments114 pages, 16 figures, 38 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑