发表机构
Amazon; National Taiwan University; Plaksha University; RAND Corporation; Algoverse; Google DeepMind(亚马逊; 国立台湾大学; 普拉卡莎大学; 兰德公司; Algoverse; 谷歌DeepMind)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究探究指令微调LLM在系统与用户指令冲突中的仲裁机制,构建含41对约束的基准评估8个模型,发现反层级模型Llama-3.1-8B的冲突结果可从残差流激活线性解码,干预效果取决于读出几何结构。
AI 中文摘要
我们研究经指令微调的大型语言模型(LLM)如何在系统指令与用户指令的直接冲突中进行仲裁。我们引入了一个包含41对约束的基准,搭配确定性验证器,并在匹配的基线、冲突及同通道控制条件下评估8个模型。行为层面,模型通过系统权威差异分为三类:尊重层级的模型将系统通道作为权威信号;反层级模型遵循系统指令的频率低于其同通道基线的预测;无效应模型则几乎无通道敏感性。Llama-3.1-8B是本研究中最强的反层级案例,在冲突试验中仅0.10的情况下遵循系统指令。我们以该行为失败案例探究用户偏好的仲裁是否反映内部冲突解决信号的缺失,答案是否定的:在Llama-3.1-8B上,冲突结果可从残差流激活中线性解码,平衡准确率达0.97,较仅用元数据的基线高出17个百分点,且Qwen2.5-7B与gpt-oss-20b上存在类似信号。用4种冲突内逻辑回归方向的第12层均值进行引导,可将实际系统遵从度从0.132提升至0.530,而主要为池化可分性选择的方向引导效果较差。因此,用户偏好的冲突解决可与可读的内部仲裁信号共存,成功的干预取决于读出的几何结构而非仅探测准确率。
英文摘要
We study how instruction-tuned LLMs arbitrate direct conflicts between system and user instructions. We introduce a benchmark of 41 paired constraints with deterministic verifiers and evaluate eight models under matched baseline, conflict, and same-channel control conditions. Behaviourally, the models split into three regimes by System Authority Delta: hierarchy-respecting models use the system channel as an authority signal, anti-hierarchy models follow the system less often than their same-channel baseline predicts, and no-effect models show little channel sensitivity. Llama-3.1-8B is the strongest anti-hierarchy case in our suite, following the system in only 0.10 of conflict trials. We use this behavioural failure case to ask whether user-preferring arbitration reflects the absence of an internal conflictresolution signal. It does not: on Llama-3.1-8B, the conflict outcome is linearly decodable from residual-stream activations at 0.97 balanced accuracy, 17 percentage points above a metadata-only baseline, with analogous signals on Qwen2.5-7B and gpt-oss-20b. Steering with a layer-12 mean of four per-conflict logistic-regression directions raises genuine system compliance from 0.132 to 0.530, while directions selected mainly for pooled separability steer poorly. User-preferring conflict resolution can therefore coexist with a readable internal arbitration signal, and successful intervention depends on the geometry of the readout rather than probe accuracy alone
CommentsPublished at the ICML 2026 Mechanistic Interpretability workshop, 16 pages (including appendix)