arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

不要过度思考,不要思考不足:面向智能体人工智能的自适应推理

Don't Overthink, Don't Underthink: Toward Adaptive Reasoning in Agentic AI

Md Jueal Mia, M. Hadi Amini

arXiv 2608.26442首次发表:更新:

AI 中文总结

该研究针对智能体AI中推理分配不当导致的过度/思考不足问题,在MATH-500和GAIA基准上验证相关失败模式,为自适应推理机制研究提供依据。

AI 中文摘要

大型语言模型(LLMs)的最新进展表明,推理时增加推理过程可提升复杂任务的性能。然而,许多现有方法依赖固定或预先分配的推理控制,如固定token预算、执行前难度估计或激活空间干预,且常在独立推理基准而非完整智能体工作流上评估。这些假设在智能体人工智能系统中可能不成立,在该系统中,推理需求会因规划、工具使用、记忆检索及智能体间交互动态变化。因此,推理可能要么过度要么不足,导致不必要的计算、延迟增加、规划漂移、工具过度使用或解决方案不完整。我们认为,下一代智能体人工智能的主要挑战不仅是语言模型应执行多少推理,而是如何根据不断变化的任务需求分配推理。我们将过度推理和思考不足表征为推理分配不当的常见失败模式,并在MATH-500和GAIA公共验证基准上对其进行评估。使用工具决策延迟、token消耗、token限制耗尽及答案正确性,我们的结果表明,被归类为过度推理的案例与更高的计算成本相关,但无相应的准确率提升,而被归类为思考不足的案例始终与不正确或不完整的解决方案相关。这些发现为智能体人工智能自适应推理机制的未来研究提供了动力。

英文摘要

Recent advances in Large Language Models (LLMs) have shown that increased inference-time reasoning can improve performance on complex tasks. However, many existing approaches rely on fixed or preallocated reasoning controls, such as fixed token budgets, pre-execution difficulty estimates, or activation-space interventions, and are often evaluated on standalone reasoning benchmarks rather than full agentic workflows. These assumptions may not hold in agentic AI systems, where reasoning requirements evolve dynamically through planning, tool use, memory retrieval, and agent-to-agent interactions. Consequently, reasoning can become either excessive or insufficient, resulting in unnecessary computation, increased latency, planning drift, excessive tool use, or incomplete solutions. We argue that a major challenge for next-generation agentic AI is not merely how much reasoning a language model should perform, but how it should allocate reasoning according to evolving task demands. We characterize over-reasoning and under-reasoning as recurring failure modes of misallocated reasoning and evaluate them on MATH-500 and the GAIA public validation benchmark. Using tool-decision latency, token consumption, token-limit exhaustion, and answer correctness, our results suggest that cases classified as over-reasoning are associated with higher computational cost without proportional accuracy gains, whereas cases classified as under-reasoning are consistently associated with incorrect or incomplete solutions. These findings motivate future research on adaptive reasoning mechanisms for agentic AI.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑