arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

STRUCTURALCOST:用于建模人类句子处理难度的受控阅读时间数据集

STRUCTURALCOST: A controlled reading time dataset for modeling human sentence processing difficulty

Nina Nusbaumer, Iria de-Dios-Flores, Corentin Bel, Christophe Pallier, Guillaume Wisniewski, Benoît Crabbé

arXiv 2610.08208首次发表:更新:

发表机构

LLF, CNRS, Université Paris Cité; COLT, Universitat Pompeu Fabra; Unicog, Neurospin, CEA; LNC2, ENS-PSL; INSERM, CNRS(巴黎西岱大学语言与法国语言学实验室,法国国家科学研究中心; 庞培法布拉大学计算语言学与语言技术实验室; 法国原子能委员会神经自旋神经影像研究中心,认知神经科学单元; 巴黎高等师范学院认知神经科学实验室; 法国国家健康与医学研究院,法国国家科学研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出STRUCTURALCOST数据集,包含475名参与者的40,800个观测,用于隔离长距离主谓依存解析成本,并发现语言模型低估了人类整合成本,为评估模型认知合理性提供数据。

AI 中文摘要

我们引入了STRUCTURALCOST,一个包含475名参与者和40,800个观测数据的自定步速阅读数据集,该数据集隔离了长距离主谓依存解析的处理成本。我们在NLP规模上复现了一个低统计功效的心理语言学发现,即人类在主动词处的阅读时间随依存长度增加而增加,这一效应由句法嵌入(而非线性距离)驱动。不同的语言模型——涵盖n-gram模型、SSM和Transformer——部分反映了这一分级难度特征,但低估了人类所承受的整合成本,且这一差距在架构和模型规模上持续存在。这表明这些模型捕捉了人类处理的预测成分,但未完全捕捉工作记忆所施加的整合成本。STRUCTURALCOST提供了推动语言模型认知合理性评估进展所需的数据。

英文摘要

We introduce STRUCTURALCOST, a self-paced reading dataset of 475 participants and 40,800 observations isolating the processing cost of long-distance subject-verb dependency resolution. We replicate a low-powered psycholinguistic finding at NLP scale, namely that human reading times at the main verb increase with dependency length, driven by syntactic embedding beyond linear distance. Different language models -- spanning n-gram models, SSMs, and transformers -- partially mirror this graded difficulty profile, yet underestimate the integration cost humans incur, with a gap that persists across architectures and model sizes. This suggests these models capture the predictive component of human processing but not the full integration cost that working memory imposes. STRUCTURALCOST provides data needed to drive progress toward evaluating the cognitive plausibility of language models.

CommentsWill be published at EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑