arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2505.21765cs.AI

不要想得更长,要想得更明智:优化大型推理模型的思维动态

Don't Think Longer, Think Wisely: Optimizing Thinking Dynamics for Large Reasoning Models

  • University of California, Los Angeles(加州大学洛杉矶分校)
  • University of Maryland, College Park(马里兰大学学院公园分校)

机构由 AI 辅助整理,请以论文原文为准。

Sohyun An, Ruochen Wang, Tianyi Zhou, Cho-Jui Hsieh

更新

AI总结:

针对大型推理模型因过度思考导致的计算浪费与性能下降问题,提出动态优化框架,通过识别并促进有利的思维模式、移除不利模式来优化推理路径,显著降低计算开销并提升准确率。

AI中文摘要:

尽管大型推理模型(LRMs)近期取得的成功通过使用强化学习优化最终答案准确率,显著提升了大型语言模型(LLMs)的推理能力,但它们也可能由于过度思考而导致输出长度急剧增加,其特征是不必要的复杂推理路径不仅浪费计算资源,还可能降低性能。我们假设这种低效源于LRMs在正确位置动态选择恰当的模块化推理策略(称为思维模式)的能力有限。为了验证这一假设,我们提出了一个动态优化框架,将模型生成的推理路径分割成不同的思维模式,系统地识别并促进能改善答案的有利模式,同时移除不利模式。实证分析证实,我们优化后的思维路径产生了更简洁但信息充足的轨迹,通过将注意力FLOPs减少高达47%来提升推理效率,同时维持了原本正确响应的准确率。此外,一部分原本错误的响应被转化为正确响应,在长度减少的情况下实现了15.6%的准确率提升。受优化思维路径带来提升的启发,我们应用了一种由对比次优和最优推理路径的成对数据集支持的偏好优化技术。在多个数学推理基准上的实验评估表明,我们的方法显著降低了计算开销,同时提升了推理准确率,实现了高达12%的准确率提升,并将token使用量从约5,000个减少到3,000个。

英文摘要:

While recent success of large reasoning models (LRMs) significantly advanced LLMs' reasoning capability by optimizing the final answer accuracy using reinforcement learning, they may also drastically increase the output length due to overthinking, characterized by unnecessarily complex reasoning paths that waste computation and potentially degrade the performance. We hypothesize that such inefficiencies stem from LRMs' limited capability to dynamically select the proper modular reasoning strategies, termed thinking patterns at the right position. To investigate this hypothesis, we propose a dynamic optimization framework that segments model-generated reasoning paths into distinct thinking patterns, systematically identifying and promoting beneficial patterns that improve the answer while removing detrimental ones. Empirical analysis confirms that our optimized thinking paths yield more concise yet sufficiently informative trajectories, enhancing reasoning efficiency by reducing attention FLOPs by up to 47% while maintaining accuracy for originally correct responses. Moreover, a non-trivial portion of originally incorrect responses are transformed into correct ones, achieving a 15.6% accuracy improvement with reduced length. Motivated by the improvement brought by the optimized thinking paths, we apply a preference optimization technique supported by a pairwise dataset contrasting suboptimal and optimal reasoning paths. Experimental evaluations across multiple mathematical reasoning benchmarks reveal that our method notably reduces computational overhead while simultaneously improving reasoning accuracy, achieving up to a 12% accuracy improvement and reducing token usage from approximately 5,000 to 3,000 tokens.

补充信息

↑