arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37312cs.LGcs.AIcs.CL

隐藏推理必然泄露,但未必可读:思维链监控的基本机遇与局限

Hidden Reasoning Must Leak, but Need Not Be Readable: Fundamental Opportunities and Limits for Chain-of-Thought Monitoring

发表机构萨尔大学 · 祖斯ELIZA学院
查看机构详情
  • Saarland University(萨尔大学)
  • Zuse School ELIZA(祖斯ELIZA学院)

机构由 AI 辅助整理,请以论文原文为准。

Mohammadali Mohammadkhani, Madhava Krishna, Yash Sarrof, Michael Hahn

首次发表
浏览论文内容

中文总结 AI 辅助

本研究探讨推理模型能否隐藏思维链中的计算,发现简单计算可隐蔽,但复杂任务必然泄露近线性信息,且该泄露在密码学假设下可能不可读,揭示了CoT监控的机遇与局限。

中文摘要 AI 辅助

推理模型能否欺骗思维链(CoT)监控器,在其思考痕迹中执行隐藏计算而不被揭示?我们表明,答案取决于底层任务难度和模型规模。简单计算可以隐蔽执行;然而,超过一个依赖于模型规模的阈值后,成功解决任务必然会在思维链中泄露关于隐藏任务输入的近线性信息量。因此,足够复杂的隐藏计算总会留下信息论足迹。然而,令人担忧的是,这种泄露未必可读:在合理的密码学假设下,即使单层Transformer也能在线加密其推理,使得任何多项式时间监控器都无法提取关于隐藏计算的信息。总体而言,我们的理论和实证结果提供了对思维链监控机遇与局限的整体视角。

英文摘要

Can reasoning models trick chain of thought (CoT) monitors and perform hidden computation without revealing it in their thinking traces? We show that the answer depends on the underlying task difficulty and the model size. Simple computations can be performed covertly; however, beyond a threshold depending on model size, successfully solving the task necessarily leaks a near-linear amount of information about the covert task input into the CoT. Therefore, sufficiently complex hidden computation always leaves an information-theoretic footprint. However, concerningly, this leakage need not be readable: Under plausible cryptographic assumptions, even a one-layer Transformer can encrypt its reasoning online so that no polynomial-time monitor can extract information about the hidden computation. Overall, our theoretical and empirical results provide a holistic view of both the opportunities and the limitations of CoT monitoring.

↑