发表机构
Georgetown University; The Ohio State University; UC San Diego(乔治城大学; 俄亥俄州立大学; 加州大学圣迭戈分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究在不同模型规模与训练语料下系统比较多种注意力记忆机制,发现对中间词元内容敏感的约束与人类阅读时间匹配度最高,且动态记忆课程下心理测量拟合度与语法能力存在分离,提示Transformer无法作为通用认知模型。
AI 中文摘要
基于Transformer的语言模型被广泛用作人类语言处理的模型,但其注意力机制可无损访问全部前文上下文,与人类有限的记忆系统不同。我们假设在Transformer的注意力机制中加入记忆约束可提升其对人类行为数据的拟合度。过往研究仅单独探索各类约束,我们则在不同模型规模与训练语料下,对多种基于注意力的记忆机制开展系统比较,同时评估其对人类阅读时间的心理测量预测能力与语法能力。我们还对比了静态约束(约束强度在整个训练过程中固定)与动态记忆课程。结果发现,对中间词元内容敏感的约束始终与人类阅读时间的匹配度最高,优于基于距离的约束。我们观察到动态记忆课程下心理测量拟合度与语法能力存在分离,表明Transformer无法作为通用的认知模型。
英文摘要
Transformer-based language models are widely used as models of human language processing, yet their attention mechanisms allow lossless access to the full preceding context, unlike the limited memory systems of humans. We hypothesize that installing memory constraints into transformers' attention mechanisms can improve their fit to human behavioral data. While previous work has explored individual constraints in isolation, we conduct a systematic comparison of multiple attention-based memory mechanisms across different model sizes and training corpora, evaluating both psychometric predictive power for human reading times and grammatical competence. We additionally compare static constraints, in which the constraint strength is fixed throughout training, to dynamic memory curricula. We find that constraints that are sensitive to the content of intervening tokens consistently achieve the highest alignment with human reading times, outperforming distance-based constraints. We observe a dissociation between psychometric fit and grammatical competence under dynamic memory curricula, suggesting that Transformers cannot serve as a one-size-fits-all cognitive model.