arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EchoCoT:从大型推理模型中提取隐藏的思维链

EchoCoT: Extracting Hidden Chain-of-Thought from Large Reasoning Models

Yiting Qu, Ziqing Yang, Chi Cui, Ye Leng, Junjie Chu, Yang Zhang

arXiv 2608.20055首次发表:更新:

发表机构

CISPA Helmholtz Center for Information Security(CISPA亥姆霍兹信息安全中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出EchoCoT多步骤攻击方法,通过API交互可从开源及前沿专有大型推理模型中近乎逐字提取隐藏思维链,该方法还具备泛化性,凸显隐藏CoT提取的安全风险。

AI 中文摘要

隐藏的思维链(Chain-of-Thought, CoT)轨迹,尤其是来自前沿专有大型推理模型(Large Reasoning Models, LRMs)的隐藏CoT,是极具价值的模型资产。然而,这些隐藏CoT能否从黑盒模型中直接提取,在很大程度上仍未被探索。本研究系统探究是否可通过API交互从黑盒LRMs中近乎逐字提取隐藏CoT。我们识别出工具调用之间此前被忽视的推理重放表面,并开发出EchoCoT,这是一种利用API返回的保真度信号迭代提取隐藏CoT的多步骤攻击方法。我们还开发了一个基于大型语言模型(LLM)的优化框架,可自动在各类数据集上搜索有效的通用注入轨迹。我们在3个开源LRMs和5个前沿专有LRMs上对EchoCoT进行评估:在开源LRMs上,EchoCoT实现了高达66.4%的近乎逐字提取成功率,提取的轨迹长度与目标长度偏差在10%以内,且至少90%的标记与目标CoT完全匹配;相同的注入轨迹还可泛化到未见过的数据集,在相同标准下实现高达80%的提取成功率。对于测试的前沿专有LRMs,大部分提取的CoT与提供商报告的推理长度及可用CoT摘要高度吻合。EchoCoT还可提取极长的CoT:在Gemini-2.5上,它从32948个标记的目标中提取了33463个标记。这些结果表明,隐藏CoT提取是一种实际存在的安全风险,凸显了更好保护隐藏CoT资产的必要性。

英文摘要

Hidden chain-of-thought (CoT) traces, especially those from frontier proprietary large reasoning models (LRMs), are valuable model assets. Yet whether these hidden CoTs can be directly extracted from black-box models remains largely unexplored. In this work, we systematically study whether hidden CoTs can be extracted near-verbatim from black-box LRMs through API interactions. We identify a previously overlooked reasoning replay surface between tool calls and develop EchoCoT, a multi-step attack that iteratively extracts hidden CoTs using API-returned fidelity signals. We further develop an LLM-based optimization framework that automatically searches for an effective universal injection trajectory across various datasets. We evaluate EchoCoT on three open-source and five frontier proprietary LRMs. On open-source LRMs, EchoCoT achieves up to 66.4\% near-verbatim extraction success, with the extracted trace length within 10\% of the target and at least 90\% of tokens exactly matching the target CoT. The same injection trajectory also generalizes to unseen datasets, achieving up to 80\% extraction success under the same criterion. For tested frontier proprietary LRMs, a substantial fraction of extracted CoTs closely align with provider-reported reasoning lengths and available CoT summaries. EchoCoT can also extract very long CoTs: on Gemini-2.5, it extracts 33,463 tokens from a 32,948-token target. These results establish hidden-CoT extraction as a practical security risk and highlight the need to better protect hidden CoT assets.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑