两个小型LLM中不存在可用的线性“投降方向”:激活引导声明的验证协议及跨家族顺从行为研究
No Usable Linear "Capitulation Direction" in Two Small LLMs: A Validation Protocol for Activation-Steering Claims, and a Cross-Family Behavioral Study of Sycophancy Under Pushback
浏览论文内容
中文总结 AI 辅助
本研究通过验证协议证明两个小型LLM中不存在可用的线性投降方向,反驳行为依赖模型且认知上有害,现有探针信号为过拟合。
中文摘要 AI 辅助
语言模型在用户反驳时经常放弃正确答案。我们在来自不同家族的两个小型指令微调模型Qwen2.5-1.5B和Llama-3.2-1B上,基于TriviaQA数据集研究这一现象:模型先作答,随后受到四种脚本化反驳风格之一的挑战,再作答一次。在初始答案正确的情况下,模型在41.8%和43.1%的回合中翻转为错误答案。哪种压力有效是模型的性质,而非压力的性质:预先指定的同一问题内配对比较(单纯质疑 vs. 情感诉求)在不同家族中呈相反方向的Bonferroni显著性(Qwen:单纯质疑 > 情感诉求,OR 2.5,p=.040;Llama:情感诉求 > 单纯质疑,OR 4.0,p=.001)。失败模式也依赖于模型:Llama放弃答案而不重新承诺的频率是Qwen的六倍(8.2% vs. 1.4%)。相同的反驳仅能修复约13%的初始错误答案;反驳在认知上净具有破坏性。接着我们探究投降是否可从响应前残差流中线性解码,这是在该位置进行引导向量干预的先决条件。一个朴素的均值差探针看似成功(样本内AUROC 0.81/0.71),但结合问题级交叉验证、打乱标签零分布和已知方向阳性对照的验证协议表明该信号是过拟合的:Qwen中最佳交叉验证AUROC为0.582,Llama中为0.548,两者均接近或低于其置换阈值,且远低于预先注册的可用性标准0.70,而相同流程在两个模型中均以AUROC 1.000恢复反驳存在性对照方向。我们进一步量化了一个测量风险:子串评分将投降低估了18-24个百分点。代码、提示、转录和分析均已发布。
英文摘要
Language models frequently abandon correct answers when users push back. We study this in two small instruction-tuned models from different families, Qwen2.5-1.5B and Llama-3.2-1B, over TriviaQA: the model answers, is challenged with one of four scripted pushback styles, and answers again. Conditioned on an initially correct answer, the models flip to a wrong answer in 41.8% and 43.1% of episodes. Which pressure works is a property of the model, not the pressure: the same within-question paired comparison (bare doubt vs. emotional appeal), specified in advance, is Bonferroni-significant in opposite directions across families (Qwen: bare doubt > emotional, OR 2.5, p=.040; Llama: emotional > bare doubt, OR 4.0, p=.001). Failure mode is also model-dependent: Llama abandons answers without recommitting at six times Qwen's rate (8.2% vs. 1.4%). Identical pushback repairs initially wrong answers only ~13% of the time; pushback is net epistemically destructive. We then ask whether capitulation is linearly decodable from the pre-response residual stream, a prerequisite for steering-vector interventions at that locus. A naive difference-in-means probe appears to succeed (in-sample AUROC 0.81/0.71), but a validation protocol combining question-level cross-validation, shuffled-label nulls, and a known-direction positive control shows the signal is overfitting: the best cross-validated AUROC is 0.582 in Qwen and 0.548 in Llama, both near or below their permutation thresholds and far under a pre-registered usability bar of 0.70, while the identical pipeline recovers a pushback-presence control direction at AUROC 1.000 in both. We further quantify a measurement hazard: substring grading underestimates capitulation by 18-24 percentage points. Code, prompts, transcripts, and analysis are released.