人类阅读与大语言模型处理中概念性与指称性干扰的不同动力学特征
Distinct dynamics of conceptual and referential disruptions in human reading and large language model processing
- Universitat Pompeu Fabra(庞培法布拉大学)
- Institut Català de Recerca i Estudis Avançats (ICREA)(加泰罗尼亚高级研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究对比人类阅读与大语言模型处理,发现概念性与指称性干扰的动力学特征存在差异,为两种意义类型的可区分处理机制提供了聚合证据。
AI中文摘要:
语言意义以概念内容为基础,当词语进入话语时,对特定实体的指称便从中产生。为探究与意义这两个维度相关的处理动力学,我们在短篇叙事中选择性干扰概念性或指称性信息,并追踪其在人类自定步速阅读以及大语言模型的预测与表征处理中产生的效应。在人类阅读中,概念性干扰会产生强烈但局部的处理成本,在出现被扭曲词语后立即显现,达到早期峰值后迅速下降;指称性干扰产生的效应较弱,会在后续词语中更缓慢地减弱,且受句子边界的调节作用更强。在大语言模型中,两种干扰均在被操纵的词语处立即显现。上下文模型惊讶值呈现出与人类阅读高度相似的模式:概念性干扰产生更大、更局部集中的效应,且衰减迅速;而指称性干扰产生更小、更渐进的下游效应。输出层表征则呈现不同模式:指称性干扰产生更大的初始位移,随后两种干扰均表现出幂律衰减。这些结果共同为两种意义类型的可区分处理动力学提供了聚合证据:概念信息施加更局部集中的整合成本,而指称性信息则参与维持话语层面同一性的更分布式过程。
英文摘要:
Linguistic meaning is grounded in conceptual content, from which reference to particular entities emerges as words enter discourse. To examine the processing dynamics associated with these two dimensions of meaning, we selectively disrupted conceptual or referential information in short narratives and traced the resulting effects in human self-paced reading and in the predictive and representational processing of large language models. In human reading, conceptual disruptions produced a strong but localized processing cost, emerging immediately after the distorted word, reaching an early maximum, and then declining rapidly. Referential disruptions produced weaker effects, which decreased more gradually across subsequent words, and were more strongly modulated by sentence boundaries. In the language model, both disruptions emerged immediately at the manipulated word. Contextual model surprisal showed a pattern closely paralleling human reading: conceptual disruption produced a larger, more locally concentrated effect that decayed rapidly, whereas referential disruption produced a smaller and more gradual downstream effect. Output-layer representations showed a different pattern: referential disruption produced a larger initial displacement, while both distortions were subsequently characterized by power-law decay. Together, these results provide convergent evidence for distinguishable processing dynamics of two types of meaning: conceptual information imposes a more locally concentrated integration cost, whereas referential information engages a more distributed process of maintaining discourse-level identity.