arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32119cs.CL

使用语言模型建模句子理解过程中语境和指代的影响

Using LMs to Model the Effects of Context and Coreference during Sentence Comprehension

Kohei Kajikawa, Lin Ai, Tatsuki Kuribayashi, Ethan Gotlieb Wilcox

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过系统变化GPT-2上下文窗口,发现扩展上下文提升心理语言学拟合,并证明追踪长距离指代关系是关键机制。

中文摘要 AI 辅助

语言模型(LMs)常被用作模拟人类语言处理的工具。近期研究表明,通过严格限制语言模型的上下文窗口,模拟人类工作记忆限制,可以提高其对人类心理语言学数据的拟合度。然而,这种严格的记忆衰减方法可能忽视了人类对长距离结构表征(如话语结构)的依赖。在本工作中,我们系统性地变化GPT-2的上下文窗口大小,在四个大规模自然主义英语阅读时间数据集上进行实验,观察到U形关系:尽管受限上下文(少于20个标记)成功捕捉了局部记忆限制,但扩展上下文(500至1000个标记)最终产生了最高的整体心理语言学拟合度。为了探究这一优势背后的机制,我们进行了一项反事实的推理时实验,通过将重复出现的话语实体代词化来破坏跨句实体链。模糊这些结构联系显著降低了较大上下文窗口的预测能力,降幅达20%至40%。我们的实验表明,追踪长距离指代关系是语言模型惊异度与人类阅读行为对齐的重要因素,并近似了人类理解者在语言处理过程中使用全局话语关系的程度。

英文摘要

Language models (LMs) are often used as a tool to model human language processing. Recent studies suggest that severely restricting LMs' context window improves their fit to human psycholinguistic data by simulating human working memory constraints. However, it is possible that this strict memory-decay approach overlooks humans' reliance on long-range structural representations, such as discourse structre. In this work, we systematically vary the context window size of GPT-2 across four large-scale naturalistic English reading-time datasets and observe a U-shaped relationship: Although restricted contexts (< 20 tokens) successfully capture local memory limitations, expanded contexts (500--1,000 tokens) ultimately yield the highest overall psycholinguistic fit. To investigate the mechanism driving this benefit, we conduct a counterfactual inference-time experiment that disrupts cross-sentential entity chains by pronominalizing repeated discourse entities. Obscuring these structural linkages significantly degrades the predictive power of larger context windows by 20% to 40%. Our experiments demonstrate that tracking long-range coreference relations is one important factor for the alignment between LM surprisal and human reading behavior, and approximate the extent to which human comprehenders use global discourse relations during language processing.

补充信息

↑