arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27756cs.CL

承载上下文:用于评估语言推理中上下文依赖程度的问题损伤分数

Load-Bearing Context: The Question Damage Score for Evaluating Context Reliance in Linguistic Reasoning

  • CUNY(纽约城市大学)

机构由 AI 辅助整理,请以论文原文为准。

Neh Majmudar, Elena Filatova

AI总结:

该研究提出基于53道英国语言学奥林匹克谜题的诊断框架,用问题损伤分数评估大型语言模型的上下文依赖,发现前沿模型在承载上下文移除后仍常产生正确答案,推动相关推理与可解释性研究。

AI中文摘要:

判断大型语言模型是从上下文还是从先验知识中推导答案,仍是一个根本性挑战。自包含的语言学奥林匹克谜题提供了一个可控场景,其中所有答案仅来自专家设计的上下文示例,无需外部知识。删除单个上下文示例会消除特定问题所需的信息,同时保持谜题其余部分不变。我们利用这一点引入了一个用于分析单个上下文示例的诊断框架。使用53道英国语言学奥林匹克谜题,我们通过删除单个上下文示例生成了两个修改变体:(1)均匀随机删除,(2)针对性删除(受纠错码启发),以删除唯一携带必要信息的结构性承载示例。我们使用问题损伤分数将谜题形式化为脆弱或稳健。在要求信息不足时弃权(不执行)的指令下评估三个前沿大型语言模型,我们发现它们很少弃权,在承载上下文被移除后仍经常继续产生正确答案。这些发现推动了对基于上下文的推理、先验知识、记忆和语言推理的进一步研究。除弃权外,该框架还支持对上下文依赖的细粒度分析,包括因果干预、停止集分析、针对性污染研究和机制可解释性。

英文摘要:

Determining whether large language models derive answers from context or prior knowledge remains a fundamental challenge. Self-contained linguistic olympiad puzzles provide a controlled setting where all answers derive solely from expert-designed context examples without external knowledge. Removing individual context examples can eliminate information needed for specific questions while leaving the rest of the puzzle unchanged. We leverage this to introduce a diagnostic framework for analyzing individual context examples. Using 53 UK Linguistics Olympiad puzzles, we generate two modified variants by deleting a single context example: (1) uniform random deletion, and (2) targeted deletion (inspired by error-correcting codes) to remove a structurally load-bearing example uniquely carrying necessary information. We formalize this impact using a Question Damage Score to classify puzzles as fragile or robust. Evaluating three frontier LLMs under instructions to abstain when information is insufficient, we find they rarely abstain, often continuing to produce correct answers after load-bearing context is removed. These findings motivate further investigation into context-based reasoning, prior knowledge, memorization, and linguistic inference. Beyond abstention, the framework enables fine-grained analyses of context reliance, including causal interventions, stopping-set analysis, targeted contamination studies, and mechanistic interpretability.

补充信息

↑