arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

过度思考:放大推理权重以提取学习到的秘密

Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets

Jack Hopkins, Dipika Khullar, Fabien Roger

arXiv 2607.08173首次发表:更新:

AI 中文总结

研究针对语言模型黑盒审计可能遗漏隐藏信息的问题,提出“过度思考”方法,通过特定模型构建及新策略放大推理,经实验证明该方法能更频繁揭示隐藏信息,不同秘密浮现方式有别。

AI 中文摘要

语言模型的黑盒审计是部署前的重要工具,但可能遗漏细微的不一致形式和隐藏信息。为在审计过程中更好地引出隐藏信息,我们引入“过度思考”,即利用推理任务向量放大推理模型出声思考倾向的过程。给定非推理指令模型\(M\)和推理提炼模型\(R\)的参数,定义过度思考模型\(\boldsymbol{\theta}_{\mathcal{O}_\alpha} = \boldsymbol{\theta}_{\mathcal{M}} + \alpha(\boldsymbol{\theta}_{\mathcal{R}} - \boldsymbol{\theta}_{\mathcal{M}})\)(\(\alpha > 1\))。还引入新的逐层衰减策略。通过四个实验设置证明过度思考模型更易揭示隐藏信息,发现推理放大揭示训练中获得的秘密或意外行为的频率比原始推理模型高\(10\)倍,秘密浮现方式取决于秘密类型。

英文摘要

Black box auditing of language models is an essential pre-deployment tool, but it may miss subtle forms of misalignment and hidden information. To better elicit hidden information during an auditing process, we introduce \emph{overthinking}: the process of using reasoning task vectors to amplify the propensity to think out loud of reasoning models. Given the parameters of a non-reasoning instruct model $M$ and reasoning-distilled model $R$, we define the \emph{overthinking model} as $\boldsymbolθ_{\mathcal{O}_α} = \boldsymbolθ_{\mathcal{M}} + α(\boldsymbolθ_{\mathcal{R}} - \boldsymbolθ_{\mathcal{M}})$, where $α> 1$ amplifies reasoning beyond the pure reasoning model $R$. Additionally, we introduce new layer-wise attenuation strategies that selectively amplify reasoning without losing quality and coherence of model outputs. We demonstrate that overthinking models are more likely to reveal hidden information across four experimental settings, across 2B-32B models. Our findings suggest that reasoning amplification may surface secrets or unintended behaviors acquired during training up to $10\times$ more frequently than the original reasoning model. How secrets surface depends on the secret type: some require perturbation along the reasoning direction, while others yield to any sufficiently large weight perturbation.

CommentsAccepted at ICML 2026. 9 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑