发表机构
Massachusetts Institute of Technology(麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究采用归因修补方法定位LLMs的算术推理电路,发现电路重叠度可预测其跨格式泛化能力,表现媲美监督探针且无需标记数据。
AI 中文摘要
在包括算术推理在内的多种推理形式中,人类能轻松实现输入格式表面变化的泛化:任何能解2+5的人也能解“two plus five”。相比之下,大型语言模型(LLMs)对提示的表面变化更脆弱:例如,它们能几乎完美地解决数字算术问题,但对相同问题的文字表述形式准确率大幅降低。本研究探究能否从模型内部预测跨格式泛化能力。研究采用归因修补(attribution patching)方法,先在英语(“two plus five”)、西班牙语(“dos más cinco”)、意大利语(“due più cinque”)三种语言中,分别定位模型解决数字算术问题(如2+5)与文字算术问题的电路;再测试模型自身数字电路的重叠度是否能预测其对文字格式的泛化能力。研究在三个层面验证了该假设:电路重叠度可解释三种文字格式的相对难度、哪些模型泛化效果最佳、哪些样本能被正确求解,其表现可与监督探针(supervised probes)相媲美,且无需标记数据。
英文摘要
In many forms of reasoning, including arithmetic reasoning, generalizing across superficial changes in input format is effortless for humans: anyone who can solve 2+5 can also solve 'two plus five'. In contrast, LLMs are more brittle to surface variations of the prompts: for example, they solve numeric arithmetic problems almost perfectly but are substantially less accurate on verbal renditions of the same problems. Here, we ask whether generalization across formats can be predicted from the models' internals. Using attribution patching, we first independently localize the circuit that each model recruits to solve numeric arithmetic problems (2+5) vs. verbal ones, in three languages: English ('two plus five'), Spanish ('dos más cinco'), and Italian ('due più cinque'); then, we test whether overlap with the model's own numeric circuit predicts its generalization to the verbal formats. Indeed, we find support for this idea at three levels: circuit overlap accounts for the relative difficulty of the three verbal formats, for which models generalize best, and for which items are solved correctly, rivaling supervised probes while requiring no labeled data.