关系先于实体:语言模型事实回忆中的延迟承诺
Relation Before Entity: Deferred Commitment in Language Model Factual Recall
AI总结:
研究发现语言模型事实回忆中关系信息先于实体信息成为生成控制因素,实体承诺被延迟,但实体信息早期已可用。
AI中文摘要:
我们探究在回忆过程中,关系类型信息(如“首都”)和实体特定信息(如“法国到巴黎”)是否在最终词元位置、相同深度处因果激活。通过使用四种互补的因果诊断方法,对四个仅解码器模型和八个提示族进行测试,我们发现一个稳健的时间不对称性:关系信息在实体信息之前成为生成控制因素。在阈值0.4下,关系起始早于实体起始10-16个测试层(占网络深度的31-44%),且在阈值0.2-0.5的所有16种模型-阈值组合中,该顺序均成立。关键的是,实体信息并非早期缺失:在早期层中,实体词元修补成功率达90-100%。相反,实体对生成的承诺被延迟:实体信息在实体词元位置可用,但仅在路由到最终词元后才成为生成控制因素。
英文摘要:
We ask whether relation-type information (e.g., capital-of) and entity-specific information (e.g., France to Paris) become causally active at the final-token position at the same depth during recall. Using four complementary causal diagnostics across four decoder-only models and eight prompt families, we find a robust temporal asymmetry: relation information becomes generation-controlling before entity information does. Relation onset precedes entity onset by 10-16 tested layers (31-44% of network depth) at threshold 0.4, with the ordering holding across all 16 model-threshold combinations for thresholds 0.2-0.5. Critically, entity information is not absent early: entity-token patching succeeds at 90-100% in early layers. Instead, entity commitment to generation is deferred: entity information is available at the entity-token position but becomes generation-controlling at the final token only after being routed there.