自动还是受控?重复启动揭示基础大语言模型、指令调优大语言模型与人类的加工差异
Automatic or Controlled? Repetition Priming Reveals Divergent Processing in Base LLMs, Instruct LLMs, and Humans
浏览论文内容
中文总结 AI 辅助
该研究通过重复启动实验对比基础大语言模型、指令调优大语言模型与人类的重复信息加工差异,发现训练后模型加工模式发生质变,且Qwen 2.5家族中该差异随规模增大而增强。
中文摘要 AI 辅助
自然语言使用中词语会不断重复,但目前仍不清楚语言模型是重新激活先前的表征,还是重新评估重复出现的词语,以及训练后是否会改变这种默认行为。我们将重复启动(Shiffrin和Schneider,1977)应用于五个模型家族的15个模型(参数规模15亿至140亿),在语义分类和完形填空两项任务中开展实验,并使用相同刺激开展匹配的人类实验。我们发现,基础模型表现出自动加工特征:呈现即时启动效应,该效应在时间滞后过程中保持稳定,部分在上下文移除后仍存在,且与对先前出现内容的注意力相关。指令调优模型表现出受控加工特征:其启动效应随时间滞后衰减,在缺少预期上下文时消失,且在更大规模时反转成干扰效应。在Qwen 2.5家族内,这种分离随模型规模单调增加,表明训练后会逐步改变重复加工的模式。人类表现出混合特征,其时间滞后敏感的启动效应类似指令调优模型,但无干扰效应,说明两种模型类型均未完全捕捉人类认知。我们的发现揭示了语言模型在训练后处理重复信息的方式发生了质性转变,并为模型行为差异提供了机制证据。
英文摘要
Words recur constantly in natural language use, yet it remains unclear whether language models reactivate prior representations or re-evaluate repeated words afresh, and whether post-training changes this default behavior. We apply repetition priming (Shiffrin and Schneider, 1977) to 15 models across five model families (1.5B-14B parameters) in two tasks, semantic categorization and cloze completion, with matched human experiments using identical stimuli. We find that base models exhibit automatic processing: they show immediate facilitation that remains stable across lags, partially survives context removal, and correlates with attention to prior occurrences. Instruct models exhibit controlled processing: their facilitation decays with lag, collapses without expected context, and reverses to interference at larger scales. Within the Qwen 2.5 family, this dissociation increases monotonically with model scale, suggesting that post-training progressively alters repetition processing. Humans show a hybrid profile, with lag-sensitive facilitation resembling instruct models but without interference, suggesting that neither model type fully captures human cognition. Our findings reveal a qualitative shift in how language models process repeated information after post-training and provide mechanistic evidence for the divergence between model behaviors.
发表机构
- University of California, Los Angeles(加利福尼亚大学洛杉矶分校)
机构由 AI 辅助整理,请以论文原文为准。