arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27590cs.CL

MWE-ECL:可恢复的长距离上下文并不总是覆盖局部词汇先验

MWE-ECL: Recoverable Long-Range Context Does Not Always Override Local Lexical Priors

Wei He, Aline Villavicencio, Rodrigo Wilkens, Zhenyun Deng

首次发表
浏览论文内容

中文总结 AI 辅助

提出MWE-ECL双语诊断工具,检验可恢复的长距离上下文是否改变多词表达的局部语义决策,发现检索成功不保证行为影响,且模型间存在异质性。

中文摘要 AI 辅助

长上下文评估通常测试模型能否恢复远处的证据,但可恢复性并不保证行为上的影响。我们检验了一个预测:一个远处的语篇锚点可以保持明确可恢复,却未能改变对熟悉多词表达的局部优先解读;此类失败应集中在模型的无锚默认与锚点冲突的情况下,而先验正确的决策则基本保持不变。我们引入了多词表达有效上下文长度(MWE-ECL),这是一个双语诊断工具,其匹配的锚点检索、无锚先验和解释提示分别衡量明确可恢复性、模型观察到的默认行为以及锚点条件下的决策。在共享的0-128K网格上的八个英语部署面板中,先验冲突项上的检索控制准确率为0.989-1.000,先验冲突覆盖率为0.806-1.000(在正确检索条件下为0.809-1.000),先验正确决策的保持率为0.977-1.000。一个在同一提示中查询检索和解释的同调用控制重现了DeepSeek V4 Pro的差距(检索为1.000,解释为0.900-0.920),表明单独调用并非其唯一解释;其他两个模型中较小或缺失的差距限制了其普遍性。对于DeepSeek V4 Flash,单独的提示拟合测试在512K和1M处保持了完美的检索,但解释较低,而干扰项一致的线索对无锚先验的移动远大于对检索的影响;跨模型的线索效应具有异质性。一个单独报告的10族中文子集显示了类似的描述性差距,但某些模型的检索不完美阻止了仅整合的归因。因此,MWE-ECL评估了明确可恢复的远处上下文是否改变了一个竞争性的局部语义决策。

英文摘要

Long-context evaluations often test whether a model can recover distant evidence, but recoverability does not guarantee behavioral influence. We test the prediction that a distant discourse anchor can remain explicitly recoverable yet fail to change the locally preferred reading of a familiar multiword expression; such failures should concentrate when the model's no-anchor default conflicts with the anchor, while prior-correct decisions remain largely preserved. We introduce Multiword Expression Effective Context Length (MWE-ECL), a bilingual diagnostic whose matched anchor-retrieval, no-anchor prior, and interpretation prompts measure explicit recoverability, model-observed defaults, and anchor-conditioned decisions, respectively. Across eight English deployment panels on a shared 0-128K grid, retrieval-control accuracy on prior-conflict items is 0.989-1.000, prior-conflict override spans 0.806-1.000 (0.809-1.000 after conditioning on correct retrieval), and preservation of prior-correct decisions remains 0.977-1.000. A same-call control querying retrieval and interpretation in one prompt reproduces the gap for DeepSeek V4 Pro (1.000 retrieval versus 0.900-0.920 interpretation), showing that separate invocations are not its sole explanation; smaller or absent gaps in the other two models bound its generality. For DeepSeek V4 Flash, separate prompt-fit tests retain perfect retrieval with lower interpretation at 512K and 1M, while foil-consistent cues shift the no-anchor prior far more than retrieval; cross-model cue effects are heterogeneous. A separately reported 10-family Chinese subset shows similar descriptive gaps, but imperfect retrieval for some models prevents an integration-only attribution. MWE-ECL therefore evaluates whether explicitly recoverable distant context changes a competing local semantic decision.

发表机构

  • University of Exeter(埃克塞特大学)
  • University of Cambridge(剑桥大学)

机构由 AI 辅助整理,请以论文原文为准。

↑