AI 中文总结
该研究提出OODA-Tool策略,通过分离状态维护与动作实现缓解状态-动作竞争,经Qwen3模型在多轮多工具任务上验证,可提升多轮工具使用的任务成功率,尤其对小模型及依赖多轮信息的任务效果显著。
AI 中文摘要
可靠的多轮工具使用要求智能体维护不断演变的任务状态,并确保每个动作与该状态保持一致。然而,直接函数调用和ReAct风格的策略在同一自回归轨迹中学习状态跟踪和动作生成,这种耦合会产生状态-动作竞争:生成下一次调用的压力会覆盖或忽略交互过程中早期积累的信息。受博伊德的观察-调整-决策-行动(Observe-Orient-Decide-Act,OODA)循环启发,我们引入了OODA-Tool,这是一种类型化闭环策略,旨在通过将状态维护与动作实现分离来缓解这种竞争。OODA-Tool并非直接从交互历史中生成动作,而是通过控制器检查的中间状态路由每个决策,确保最终输出基于当前任务状态。具体而言,观察(Observe)重构任务状态,调整(Orient)确定是否需要执行,决策(Decide)形成可接受的动作结构,行动(Act)实现外部输出。我们使用Qwen3模型(参数规模从0.6B到14B)在多轮、多工具和信息不完整的设置下,将OODA-Tool与直接函数调用和ReAct策略进行评估。OODA-Tool在所有模型规模下均持续提升任务成功率,在较小模型以及动作高度依赖多轮交互和先前工具结果中积累信息的任务上,提升幅度更大。受控变体、阶段级消融实验和迁移评估进一步证明了这些改进的鲁棒性。
英文摘要
Reliable multi-turn tool use requires an agent to preserve an evolving task state and ensure that each action remains consistent with it. However, direct function-calling and ReAct-style policies learn state tracking and action generation within the same autoregressive trajectory. This coupling creates state-action competition: the pressure to produce the next call can overwrite or ignore information accumulated earlier in the interaction. Inspired by Boyd's Observe-Orient-Decide-Act cycle, we introduce OODA-Tool, a typed closed-loop policy designed to mitigate this competition by separating state preservation from action realization. Rather than generating an action directly from the interaction history, OODA-Tool routes each decision through controller-checked intermediate states, ensuring that the final output remains grounded in the current task state. Specifically, Observe reconstructs the task state, Orient determines whether execution is warranted, Decide forms an admissible action structure, and Act realizes the external output. We evaluate OODA-Tool against direct function-calling and ReAct policies using Qwen3 models ranging from 0.6B to 14B across multi-turn, multi-tool, and incomplete-information settings. OODA-Tool consistently improves task success across model sizes, with larger gains on smaller models and on tasks whose actions depend strongly on information accumulated across turns and prior tool results. Controlled variants, stage-level ablations, and transfer evaluations further demonstrate the robustness of these improvements.