基于大语言模型(LLM)的智能体在复杂任务上展现出强大能力,它们通常在整个交互轨迹的每一轮动作前执行推理。然而,并非每一轮都需要推理,因为早期生成的推理可继续为后续动作提供支持。因此,关键挑战在于判断现有推理何时仍足够、何时需要新的推理步骤,且无需依赖成本高昂的生成式验证。我们发现,移除额外推理后后续参考动作的概率下降情况,与给定早期推理下这些动作是否仍可恢复密切相关,这为估计跨轮次动作支持提供了有效且轻量的信号。基于此观察,我们提出通过跨轮次估计实现推理自适应的训练方法RACE(Reasoning Adaptation through Cross-Turn Estimation)。RACE引入似然引导的渐进推理覆盖检测(LoGiC)流程,逐步识别移除后对当前及后续参考动作影响有限的推理轮次。所得移除信号被整合到监督微调与智能体强化学习中,使策略能学习何时推理、何时直接行动。在四个代表性智能体基准上的大量实验表明,RACE大幅降低了推理成本,同时维持或提升了任务性能。
英文摘要
Large language model (LLM)-based agents have demonstrated strong capabilities on complex tasks. They typically perform reasoning before each action throughout an interaction trajectory. However, reasoning may not be necessary at every turn, as reasoning produced earlier can continue to support subsequent actions. A key challenge is therefore to determine when existing reasoning remains sufficient and when a new reasoning step is needed, without relying on costly generation-based verification. We find that decreases in the likelihood of subsequent reference actions after removing additional reasoning closely track whether those actions remain recoverable given earlier reasoning, providing an effective and lightweight signal for estimating cross-turn action support. Based on this observation, we propose Reasoning Adaptation through Cross-Turn Estimation (RACE), a training approach for adaptive agent reasoning. RACE introduces a Likelihood-Guided Progressive Reasoning Cover Detection (LoGiC) procedure that progressively identifies reasoning turns whose removal has limited impact on the current and subsequent reference actions. The resulting removal signals are incorporated into both supervised fine-tuning and agentic reinforcement learning, enabling the policy to learn when to reason and when to act directly. Extensive experiments on four representative agent benchmarks show that RACE substantially reduces reasoning cost while maintaining or improving task performance.