arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

高效推理中的Token价值不平等性研究

On the Token Value Inequality in Efficient Reasoning

Runjia Zeng, Hang Hua, Yiyang Liu, Zhiqiang Tao, Ruixiang Tang, Qifan Wang, Cheng Han, Dongfang Liu

arXiv 2609.33970首次发表:更新:

发表机构

Purdue University; Rochester Institute of Technology; MIT-IBM Watson AI Lab; University of Missouri-Kansas City; Rutgers University; Meta AI(普渡大学; 罗切斯特理工学院; MIT-IBM沃森人工智能实验室; 密苏里大学堪萨斯城分校; 罗格斯大学; Meta AI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出TokenProbe框架,基于Token级对数概率识别核心与冗余Token,通过选择性压缩冗余Token实现帕累托改进,在保持推理质量的同时将Token使用量减少至基线的76%。

AI 中文摘要

思维链推理使大语言模型在复杂任务上取得了显著的性能提升。然而,这些提升是以大幅增加Token消耗为代价的。这引出了一个基本问题:推理轨迹中的每个Token是否具有同等价值?我们提出了一个基于关键实证发现的诊断与优化框架:思维链推理序列中Token的价值高度不均匀,且这种不均匀性可以通过Token级别的对数概率信号有效刻画。我们证明,归一化的对数概率有助于区分核心Token(承载结构性和决定性推理内容)与冗余Token(探索性、低置信度的填充内容,对最终答案的直接贡献较小)。基于这些发现,我们构建了TokenProbe框架,该框架围绕两个实证发现和一个主张展开:发现首先识别Token价值不平等性,然后确立TokenProbe作为核心Token的代理;主张引入了一个高效的GRPO目标,假设选择性地压缩冗余Token可以在准确率-Token效率空间中实现帕累托改进。实验上,我们的方法在保持推理质量的同时,将Token使用量减少至基线的76%。在匹配的推理长度预算下,我们表明它甚至能超越如Gemini-3.1-Pro等强大的旗舰基线模型。主页:此https URL。

英文摘要

Chain-of-Thought reasoning has enabled large language models to achieve substantial performance gains on complex tasks. However, these gains come at the cost of dramatically increased token consumption. This raises a fundamental question: is every token in the reasoning trace equally valuable? We present a diagnostic and optimization framework grounded in a key empirical finding: the value of tokens within a CoT reasoning sequence is highly non-uniform, and this non-uniformity can be effectively characterized by token-level log probability signals. We show that normalized log probability helps distinguish core tokens, which carry structural and decisive reasoning content, from redundant tokens, which are exploratory, low-confidence filler that contributes less directly to the final answer. Building on these findings, we formulate the TokenProbe framework around two empirical findings and one claim: findings identify token value inequality first and then establish TokenProbe as a core-token proxy, and the claim introduces an efficient GRPO objective positing that selectively compressing redundant tokens can yield Pareto improvements in the accuracy-token efficiency space. Empirically, our method preserves reasoning quality while reducing the token usage by 76% of the baseline. Under matched reasoning-length budgets, we show that it can even outperform strong flagship baselines like Gemini-3.1-Pro. Homepage: https://runjia.tech/tokenprobe/.

CommentsNeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑