发表机构
BNP Paribas; Sorbonne Université; Criteo AI Lab(法国巴黎银行; 索邦大学; Criteo AI实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出MemoReason基准,通过配对事实与虚构推理任务,发现LLMs在虚构场景下性能显著下降(最高15.7%),表明参数记忆影响推理,但非直接捷径,而是更复杂机制。
AI 中文摘要
大型语言模型(LLMs)在推理基准测试中表现良好,但目前尚不清楚这是否反映了真正的上下文推理,还是依赖于其参数中记忆的事实。我们通过区分两种可能性来研究这一点:一种广泛的“记忆偏差”,即熟悉的内容能提高推理性能;以及“强参数捷径假说”,即模型完全跳过推理并直接回忆存储的答案。为了测试这些效应,我们引入了MemoReason,这是一个人工整理的基准,将事实推理任务与结构相同的虚构版本配对,其中真实实体(如人物、公司或日期)被系统地替换为同类型的虚构实体。这一设计保留了任务结构和指定的推理操作,同时改变了上下文的熟悉度,从而能够受控地测量参数记忆如何影响推理。我们对近期LLMs的评估显示,在虚构场景中,性能出现了一致且统计上显著的下降,最高达15.7%,这清楚地证明了记忆偏差的存在。然而,对虚构场景中失败问题的针对性分析表明,模型很少给出相应的事实答案,这表明直接的参数捷径并非主要失败模式。这些发现表明,参数记忆通过比简单事实回忆更复杂的机制影响推理。MemoReason为研究这些机制以及将配对事实-虚构评估扩展到更广泛的推理设置提供了一个受控框架。
英文摘要
Large Language Models (LLMs) perform well on reasoning benchmarks, but it remains unclear whether this reflects genuine contextual reasoning or reliance on facts memorized in their parameters. We investigate this by distinguishing two possibilities: a broad \textit{memorization bias}, where familiar content improves reasoning performance, and the \textit{Strong Parametric Shortcut Hypothesis}, where models skip reasoning entirely and recall stored answers. To test these effects, we introduce \textbf{MemoReason}, a human-curated benchmark that pairs factual reasoning tasks with structurally identical \fictitiousterm{} versions where real entities like people, companies, or dates are systematically replaced by \fictitiousterm{} ones of the same type. This \scorerevision{preserves task structure and specified reasoning operations} while varying the familiarity of the context, allowing controlled measurement of how the parametric memory affects reasoning. \revision{Our evaluation of recent LLMs reveals consistent and statistically significant performance drops of up to 15.7\% in the fictitious setting, demonstrating a clear memorization bias.} However, a targeted analysis of \revision{questions failed in the fictitious setting} shows that models rarely respond with the corresponding factual answer, indicating that direct parametric shortcuts are not the dominant failure mode. These findings suggest that parametric memory influences reasoning through mechanisms more complex than simple factual recall. \textbf{MemoReason} provides a controlled framework for studying these mechanisms and for extending paired factual-fictitious{} evaluation to broader reasoning settings.
CommentsPreprint