arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LUMOS:从训练数据到行为输出追踪LLMs中的参数化知识

LUMOS: Tracing Parametric Knowledge from Training Data to Behavioral Outputs in LLMs

Seoyeon Ye, Gayoung Kim, Jiyoung Hong, Sookyung Kim, Hyunsoo Cho

arXiv 2610.02902首次发表:更新:

发表机构

Ewha Womans University(梨花女子大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出LUMOS框架,利用OLMo 2透明训练语料追踪知识因果链,发现模型内部编码与行为表达存在差距且自我反思在未见内容上失效,主张将训练数据轴纳入LLM知识评估。

AI 中文摘要

当前对LLMs参数化知识的分析大多以输出为中心,仅根据模型能回答什么来推断其知识,而不验证其实际训练数据。这导致诸如正确回答反映的是真正的泛化还是死记硬背等基本问题,停留在推测而非证据层面。为解决这些模糊性,我们提出了LUMOS,一个诊断框架,利用具有完全透明训练语料库的OLMo 2,沿着从训练数据暴露到行为输出的因果链追踪知识。通过基于已验证的暴露进行 grounding,我们发现模型内部以高可分性(84%)编码稀有事实,但在行为上却未能表达(54%),尽管这种检索差距随规模增大而缩小。此外,当模型被要求对自己的答案进行自我反思时,它们在已训练内容上表现可靠(83%),但在未见内容上却降至随机基线水平(49%)。这种崩溃即使在思维链提示下也持续存在,思维链只会夸大置信度信号而非改善校准。总的来说,这些发现表明,将训练数据轴纳入LLM评估,可将推测性诊断转化为可验证的主张,我们主张该轴应成为LLM知识评估的标准组成部分。

英文摘要

Current analyses of LLMs' parametric knowledge are largely output-centric, drawing conclusions about what a model knows without verifying what it was actually trained on. This leaves fundamental questions, such as whether a correct response reflects genuine generalization or rote memorization, grounded in speculation rather than evidence. To resolve these ambiguities, we introduce LUMOS, a diagnostic framework that traces knowledge along the causal chain from training-data exposure to behavioral output, leveraging OLMo 2 with its fully transparent training corpus. By grounding analysis in verified exposure, we reveal that models internally encode rare facts with high separability (84%) yet fail to express them behaviorally (54%), though this retrieval gap narrows with scale. Furthermore, when models are asked to self-reflect on their own answers, they perform reliably on trained content (83%) but drop to random-baseline levels (49%) on unseen content. This collapse persists even under chain-of-thought prompting, which inflates confidence signals rather than improving calibration. Collectively, these findings demonstrate that incorporating the training-data axis into LLM evaluation transforms speculative diagnoses into verifiable claims, and we advocate that this axis should be a standard component of knowledge assessment in LLMs.

CommentsAccepted to NeurIPS 2026 (Poster)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑