编码但未解码:LLM句法中三级差距的分层证据
Encoded but Not Decoded: Layer-Localized Evidence for a Three-Level Gap in LLM Syntax
浏览论文内容
中文总结 AI 辅助
本文提出三级评估框架,发现LLM在句法任务中编码能力优于实际部署,差距源于解码偏好表面捷径,且该差距具有分层定位特征。
中文摘要 AI 辅助
语言模型可能以两种截然不同的方式在句法测试中失败:一是未能编码相关结构,二是虽然编码了该结构但在输出时未能使用它。仅凭行为评估无法区分这两种情况。我们提出了一个三级评估框架(行为部署、LM头读出和探针可恢复性),在相同的二元决策下对相同的测试项进行测量。使用一个紧凑的三语(英语、中文、德语)控制依赖基准,我们发现,在七个模型和所有三种语言的总体结果中,探针可恢复性优于或等于LM头读出,而LM头读出又优于或等于行为部署。在所有14个(模型、任务)条件下,可恢复性盈余从未为负。这种脱节集中在主语控制中,其中最近名词启发式给出了错误答案。最大的单一差距(0.653)出现在Qwen3-0.6B Instruct的问答任务中。该差距在Qwen3-14B Instruct中依然存在。指令微调在百分比上对部署的损害大于对编码的损害。我们排除了选项位置偏差、后期层擦除、输出格式伪影和探针训练方差等解释。该模式与偏向表面捷径的解码一致,行为-探针差距衡量了这种偏好的强度。激活修补显示该差距是分层定位的。在指令微调下,LM头解码层比探针解码层大约晚十层。这些发现表明,行为评估低估了模型编码的内容,而仅靠探针则高估了模型部署的内容。
英文摘要
A language model can fail a syntactic test in two distinct ways: by not encoding the relevant structure, or by encoding it but failing to use it at the output. Behavioral evaluation alone cannot tell these apart. We propose a three-level evaluation framework (behavioral deployment, LM-head readout, and probe recoverability) measured on the same items under the same binary decision. Using a compact trilingual (English, Chinese, German) control-dependency benchmark, we find that probe recoverability exceeds or equals LM-head readout, which in turn exceeds or equals behavioral deployment, across seven models and all three languages in the aggregate. The recoverability surplus is never negative across all 14 (model, task) conditions. The disconnect concentrates in subject-control, where a nearest-noun heuristic gives the wrong answer. The single largest gap (0.653) appears on Qwen3-0.6B Instruct in question answering. The gap persists at Qwen3-14B Instruct. Instruction tuning degrades deployment more than encoding in percentage terms. We rule out option-position bias, late-layer erasure, output-formatting artifacts, and probe-training variance. The pattern is consistent with decoding that favors surface shortcuts, and the behavior-probe gap measures the strength of that preference. Activation patching shows the gap is layer-localized. Under instruction tuning, the LM-head-decoded layer shifts approximately ten layers later than the probe-decoded layer. These findings argue that behavioral evaluation understates what models encode, while probing alone overstates what they deploy.
发表机构
- College of International Studies(国际研究学院)
- National University of Defense Technology(国防科技大学)
机构由 AI 辅助整理,请以论文原文为准。