发表机构
University of Cambridge(剑桥大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出CRT诊断框架,发现LRS仅注入全局标签偏差而非修复推理回路,为区分真正推理修复与表面标签覆盖提供了严谨方法。
AI 中文摘要
局部表征引导(Localized Representation Steering, LRS)被广泛用于修正大型语言模型中的推理缺陷。然而,标准基准评估极易被表面的标签覆盖所误导,从而产生推理回路修复的错误印象。本研究提出跨规则迁移(Cross-Rule Transfer, CRT)诊断框架,通过在模型天生具备能力的规则族上评估表征干预,对其进行审计。针对广泛存在的逻辑缺陷——矛盾盲性,评估深层LRS显示,该干预仅注入了全局标签偏差:将引导向量应用于模型已正确处理的规则(基线准确率99.6%),会因强制生成错误的矛盾预测而使性能降至40.4%。我们通过四个互补对照(直接对数几率偏差等价性、对照向量标签翻转、跨模型嫁接及浅层引导检查)支持该诊断,提供了区分真正推理修复与表面标签覆盖的严谨方法。
英文摘要
Localized Representation Steering (LRS) is widely used to correct reasoning pathologies in large language models. However, standard benchmark evaluations can easily be fooled by superficial label overrides, creating a false impression of reasoning circuit repairs. In this work, we propose Cross-Rule Transfer (CRT), a diagnostic framework that audits representational interventions by evaluating them on rule families where the model is natively competent. Evaluating late-layer LRS for a widespread logical failure, contradiction blindness, reveals that the intervention merely injects a global label bias: applying the steering vector to rules the model already handles correctly (99.6% baseline) degrades performance to 40.4% by forcing false contradiction predictions. We support this diagnosis with four complementary controls (direct logit bias equivalence, control vector label-flipping, cross-model grafting, and early-layer steering checks), providing a rigorous methodology to distinguish genuine reasoning repairs from superficial label overrides.