arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

推理前先承诺:开放权重语言模型中答案预承诺的行为再现及初步激活水平证据

Committed Before Reasoning: Behavioral Reproduction and Preliminary Activation-Level Evidence of Answer Pre-Commitment in an Open-Weight LLM

Heejin Jo

arXiv 2607.16451首次发表:更新:

AI 中文总结

研究聊天模型在简单问题上先承诺答案而非推导的现象,通过行为再现和初步激活水平证据揭示错误承诺情况,还指出方法学上问题措辞对结果的影响,强调阳性对照对解读阴性预言机结果的重要性。

AI 中文摘要

聊天模型有时会先承诺一个答案,然后给出推理来证明它,而不是推导答案,即便答案与任务前提相矛盾。研究一个简单问题:“我想洗车。洗车行在100米外。我该走路还是开车去?”答案只能是开车,但模型大多推荐走路。行为再现方面:在Qwen3 - 8B的五种系统提示条件下(210次展开),每种条件下85% - 100%的抽样展开以及100%的贪心展开中,无论思考模式与否,都会出现错误承诺,4096令牌的思考预算也无法修复。初步激活水平证据方面:在答案文本发出前的位置,用预训练、无任务特定探针训练的激活预言机探测隐藏状态,“走路”读数超过中性上下文基线(68%对17%;走路承诺展开p = 0.005,开车承诺展开p = 0.005,费舍尔精确检验)。方法学方面:固定预言机、激活和位置,仅问题措辞就能使一个阳性对照从2/16(开放式问题)变为11/16(封闭式问题);没有逐措辞的阳性对照,阴性预言机结果无法解释。

英文摘要

Chat models sometimes commit to an answer and then produce reasoning that justifies it rather than deriving it -- even when the answer contradicts a task premise. We study a minimal probe: "I want to wash my car. The car wash is 100 meters away. Should I walk or drive?" Only drive works (the car must be at the car wash), yet models overwhelmingly recommend walking. (1) Behavioral reproduction: on Qwen3-8B across five system-prompt conditions (210 rollouts), the wrong commitment occurs in 85-100% of sampled rollouts per condition and 100% of greedy rollouts, in both thinking and non-thinking modes; a 4,096-token thinking budget does not repair it. (2) Preliminary activation-level evidence: probing hidden states with a pretrained, training-free activation oracle (no task-specific probe training) at positions before the answer text is emitted, "walk" read-outs exceed a neutral-context baseline (68% vs. 17%; walk-committing rollouts p=.005, drive-committing rollouts p=.005, Fisher exact) -- notably, rollouts that eventually answer drive also read as walk-leaning before commitment (5/6). The oracle's default on unrelated content is "drive" (83%), so the read-outs are not lexical bias; stratifying by literal walk/drive occurrence shows they are not text recovery either (spans containing "drive" still read out walk; in balanced lexical fields, per-rollout walk-majorities beat a per-prompt neutral baseline 15/22 vs. 1/8, p=.01; drive-committing rollouts 6/6, p=.002). Samples are small and the within-rollout positional gradient is not significant (p=.34); we frame these results as preliminary. (3) Methodological: with fixed oracle, activations, and positions, question wording alone moves a positive control from 2/16 (open question) to 11/16 (closed); negative oracle results are uninterpretable without per-wording positive controls.

Comments8 pages. Code, data, and all reported statistics: https://github.com/JO-HEEJIN/interview_mate/tree/docs/car-wash-repro/car_wash/paper_4

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑