发表机构
University of Southern California; Arizona State University(南加州大学; 亚利桑那州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究发现大语言模型会回答结构上无法解答的问题是因识别到不可行性的信号与安全拒绝通路未对齐,而非缺失识别能力。
AI 中文摘要
大语言模型常回答结构上无法解答的问题,例如计算cot(-540°)或评估(1).startswith("1"),而非弃权(不执行)。我们探究该失败是源于缺失识别能力,还是识别到弃权的路由失效。在参数规模从17亿到70亿的指令微调模型中,隐藏状态中的单一线性方向可区分可解答与结构上不可能的数学、代码提示,表明模型在生成前就已表征出不可行性。然而该识别方向与介导训练后有害内容拒绝的规范安全-拒绝方向几乎正交。领域内基于行为定义的无效性感知方向更接近识别方向,但仅部分对齐,且仍与安全拒绝接近正交。沿识别方向进行生成时的引导,会在结构数学和代码任务上双向、剂量响应地改变无效性感知行为,而随机方向则无此效果。基础模型/指令微调模型的对比进一步显示,这种低余弦几何关系在预训练终点就已存在。因此,对不可能问题的自信回答失败,更宜解释为路由失败而非编码失败:模型存在可用的"无可行答案"信号,但安全-拒绝通路未对齐以使用该信号。
英文摘要
Large language models often answer structurally unanswerable questions, such as computing cot(-540°) or evaluating (1).startswith("1"), instead of abstaining. We ask whether this failure reflects missing recognition or failed routing from recognition to abstention. Across instruction-tuned models from 1.7B to 70B parameters, a single linear direction in the hidden state separates answerable from structurally impossible math and code prompts, showing that models represent impossibility before generation. Yet this recognition direction is nearly orthogonal to the canonical safety-refusal direction that mediates trained harmful-content refusal. An in-domain behavior-defined invalidity-aware direction is closer to recognition, but only partially aligned with it, and remains near-orthogonal to safety refusal. Generation-time steering along the recognition direction changes invalidity-aware behavior bidirectionally and dose-responsively on structural math and code cells, while random directions do not. Base/instruct comparisons further show that the low-cosine geometry is already present at the pretraining endpoint. The confident-on-impossible failure is therefore better explained as a routing failure than as an encoding failure: the model has a usable "no admissible answer" signal, but the safety-refusal pathway is not aligned to use it.
CommentsAccepted to EMNLP 2026 Main Conference. 29 pages, 6 figures