发表机构
Brown University; University of Michigan; Weinberg Institute for Cognitive Science(布朗大学; 密歇根大学; 温伯格认知科学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过言语冲突任务,探究LLM中一致性效应的机制,发现其源于权重内默认映射与上下文内规则映射的竞争,为分析此类响应竞争提供了模型系统。
AI 中文摘要
一致性效应在Stroop、flanker等冲突任务中已在心理学和神经科学中被研究了近一个世纪,但其机制基础尚未完全明晰。我们提出了一个仅含言语内容的LLM冲突任务,其中提示词词干会引出默认的同色补全,而明确的规则要么与该补全一致(一致性条件),要么与之冲突(不一致性条件)。Gemma-2-2B以及6个参数规模从410M到12B的Pythia模型均表现出强烈的默认同色倾向,且7个模型中有6个表现出强烈的一致性效应。通过因果归因分析、注意力分析和注意力消融实验,我们在这些LLM中识别出了不同的处理通路:一条涉及对表层颜色线索的短程注意力,在一致性条件下优先激活;另一条涉及对规则前缀的长程注意力,在不一致性条件下优先激活。增强默认同色倾向的微调对不同任务条件产生了不同影响,降低了不一致性任务的表现,同时提升了一致性任务的表现。相比之下,增加规则集规模会选择性损害不一致性任务的表现。这些一致的发现支持了如下解释:本任务中的一致性效应源于权重内默认映射与上下文内基于规则的映射之间的竞争。更广泛地说,我们的发现表明LLM可作为模型系统,用于分析单个学习网络内默认响应倾向与规则支配响应倾向之间竞争的机制。
英文摘要
Congruency effects, observed in conflict tasks such as Stroop and flanker tasks, have been investigated for nearly a century in psychology and neuroscience, but their mechanistic basis is not fully understood. We introduce a verbal-only LLM conflict task in which a prompt stem elicits a default same-color completion and an explicit rule either agrees with (congruent condition) or conflicts with (incongruent condition) the completion. Gemma-2-2B and six Pythia models ranging from 410M to 12B parameters showed strong default same-color tendencies, and six of seven models showed strong congruency effects. Using causal attribution analysis, attention analysis, and attention ablations, we identified distinct processing pathways in these LLMs: a pathway involving short-range attention to a superficial color cue that is preferentially activated in the congruent condition, and a pathway involving long-range attention to the rule prefix that is preferentially activated in the incongruent condition. Fine-tuning that strengthened the default same-color tendency had divergent effects on task conditions, reducing incongruent performance while increasing congruent performance. In contrast, increasing rule set size selectively impaired incongruent performance. These converging findings support an account in which congruency effects in this task arise from competition between an in-weight default mapping and an in-context rule-based mapping. More broadly, our findings illustrate how LLMs can serve as model systems for mechanistic analysis of competition between default and rule-governed response tendencies within a single learned network.
Comments23 pages, 8 figures