arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AdaThinking-E:用于自适应思考的单token熵调控

AdaThinking-E: One-Token Entropy Regulation for Adaptive Thinking

Zining Wang, Tongkun Guan, Boming Chen, Zhentao Guo, Jianqiang Liu, Chao Jin, Chen Duan, Kai Zhou, Pengfei Yan, Wei Shen, Xiaokang Yang

arXiv 2608.26141首次发表:更新:

发表机构

Meituan; MoE Key Lab of Artificial Intelligence; AI Institute; School of Computer Science, Shanghai Jiao Tong University; MAIS&NLPR, Institute of Automation, Chinese Academy of Sciences(美团; MoE人工智能重点实验室; 人工智能研究院; 上海交通大学计算机科学学院; 中国科学院自动化研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出AdaThinking-E强化学习框架,通过单token熵调控实现自适应思考,使模型在不同复杂度文档任务中兼顾准确率与效率

AI 中文摘要

多模态大语言模型通过引入显式思考过程展现出强大的文档推理能力,虽然该能力显著提升了挑战性任务的性能,但当前模型将这种深度思考统一应用于所有问题,导致简单任务出现不必要的计算开销,这不仅降低了用户体验,还对基准数据集上的准确率产生负面影响。我们明确了对自适应思考机制的关键需求,该机制可根据问题复杂度智能决定何时启动推理。为解决这一问题,我们提出AdaThinking-E,一种通过单token熵调控学习自适应思考的新型强化学习框架。我们的核心见解是,模型对是否启动思考的决策置信度可通过关键决策token处预测概率分布的熵分析来量化。这一观察推动了我们的熵调控奖励机制:训练过程自然从高熵探索(模型尝试不同的思考策略)过渡到低熵收敛(形成自信、可泛化的决策策略)。至关重要的是,该方法使模型能够内在地发现何时进行思考,无需人工干预或外部难度标签。大量实验表明,我们的方法使模型在各类文档任务中,既对复杂问题准确,又对简单问题高效。

英文摘要

Multimodal large language models have demonstrated strong document reasoning capabilities by incorporating explicit thinking processes. While this capability significantly improves performance on challenging tasks, current models apply such deep reasoning uniformly to all questions, resulting in unnecessary computational overhead for simple task. This not only degrades user experience but also negatively impact accuracy on benchmark datasets. We identify the critical need for adaptive thinking mechanisms that can intelligently determine when to engage reasoning based on question complexity. To address this, we propose AdaThinking-E, a novel reinforcement learning framework that learns adaptive thinking through one-token entropy regulation. Our key insight is that model confidence in the decision to engage thinking (or not) can be quantified through entropy analysis of the predicted probability distribution at critical decision tokens. This observation motivates our entropy-governed reward mechanism: the training process naturally transitions from high-entropy exploration, where the model experiments with different thinking strategies, to low-entropy convergence with confident, generalizable decision-making policies. Crucially, this approach enables models to intrinsically discover when to think without requiring manual intervention or external difficulty labels. Extensive experiments demonstrate that our approach enables models to be both accurate on complex problems and efficient on simple ones across diverse document tasks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑