硬决策层:Transformer中确定推理的证据
The Hard Decision Layer: Evidence for Committed Inference in Transformers
浏览论文内容
中文总结 AI 辅助
研究基于Transformer的语言模型在多选型问答中做预测的位置与方式,识别出硬决策层,经多模型和数据集验证其在无学习路由策略时一致出现且对微调不变,HDL处准确率显著提升,为Transformer推理提供见解与机会。
中文摘要 AI 辅助
我们研究了基于Transformer的语言模型在多项选择题回答中做出预测的位置和方式。我们识别出了“硬决策层”(HDL),这是一种自然的架构属性,在推理过程中答案选项排名会突然稳定下来。通过对四个语言模型(Qwen、Llama、Granite、Mistral)和四个基准数据集的实证验证,表明在没有学习路由策略的情况下HDL会一致出现。我们还表明HDL对微调不变。结果显示在HDL处准确率显著提高:高达+0.61(Qwen在常识问答上),之后性能稳定。对标签格式和问题复杂度的系统消融证实了该现象对模型架构至关重要。这些发现为Transformer推理提供了机制性见解,并为高效推理和模型引导提供了机会。
英文摘要
We investigate where and how transformer-based language models commit to predictions in multiple-choice question answering. We identify the _Hard Decision Layer_ (HDL), a natural architectural property where answer option rankings stabilize abruptly during inference. Empirical validation across four language models (Qwen, Llama, Granite, Mistral) and four benchmark datasets demonstrates consistent HDL emergence without learned routing policies. We also show that the HDL is invariant to fine-tuning. Our results reveal striking accuracy improvements at the HDL: up to +0.61 (Qwen on CommonsenseQA), after which performance stabilizes. Systematic ablations on label formats and problem complexity confirm the phenomenon is fundamental to model architecture. These findings offer mechanistic insights into transformer inference and suggest opportunities for efficient reasoning and model steering. All code and results required to reproduce this work are available in https://github.com/Mystic-Slice/hard-decision-layer
发表机构
- University of Southern California(南加州大学)
- Information Sciences Institute(信息科学研究所)
机构由 AI 辅助整理,请以论文原文为准。