arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26637cs.CLcs.AIcs.CR

能力强但简洁:提取并刻画前沿模型中的隐藏思维链

Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models

Xiaoyu Luo, Tao Ren, Wenrui Yu, Xiao Li, Qiongxiu Li, Johannes Bjerva

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过标准API诱导闭源前沿模型外部化隐藏思维链,验证其与原生推理性能相当,并揭示Astra采用令牌高效的定向推理,为理解前沿模型推理提供行为视角。

中文摘要 AI 辅助

前沿语言模型能力的快速提升被广泛归因于推理能力的改进,然而,由于闭源系统中的原始思维链(CoT)轨迹被隐藏,这一点无法得到验证。通过标准API功能注册一个简单的自定义工具,我们诱导前沿模型将中间推理过程外部化。由于这些轨迹可能反映事后合理化而非真正的推理,我们首先在开源模型上将其与原生思维链进行对比评估,并将评估扩展到包括GPT-6 Astra在内的闭源前沿模型。我们发现,在竞赛数学、科学和代码生成任务中,提取的推理与原生推理性能相匹配,且显著优于无推理基线。随后,我们刻画了前沿模型如何组织其中间推理。在令牌效率、推理步骤类型和诱导推理树方面,我们识别出模型在外部化、压缩和组织推理方面的系统性差异。我们发现,Astra表现出令牌高效的定向推理,更早地选择正确轨迹,同时在内部解决基本步骤,仅外部化关键推理。这些发现为超越基准分数理解前沿模型推理提供了行为视角。

英文摘要

The rapid capability gains of frontier language models are widely attributed to improved reasoning abilities, yet this cannot be verified as raw CoT traces in closed-source systems are hidden. By registering a simple custom tool through a standard API feature, we induce frontier models to externalize intermediate reasoning. Because these traces may reflect post-hoc rationalization rather than genuine reasoning, we first evaluate against native CoT on open-source models and extend to closed-source frontier models including GPT-6 Astra. We find that the extracted reasoning matches native reasoning performance and substantially outperforms no-reasoning baselines, across competition mathematics, science, and code generation. We then characterize how frontier models structure their intermediate reasoning. Across token efficiency, reasoning-step types, and induced reasoning trees, we identify systematic differences in how models externalize, compress, and organize reasoning. We find that Astra exhibits token-efficient directed reasoning, selecting a correct trajectory earlier, while resolving elementary steps internally and externalizing only crucial reasoning. These findings provide a behavioral lens on frontier-model reasoning beyond benchmark scores.

发表机构

  • Aalborg University(奥尔堡大学)
  • Seafill

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑