发表机构
Fudan University; Honor Device Co Ltd.; Peking University(复旦大学; 荣耀设备有限公司; 北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对动态意图波动下长程工具调用的意图偏差与无限API循环问题,提出IACM-RL框架,结合DynamicIntent流水线与BeliefState上下文管理器,在多基准上实现性能提升。
AI 中文摘要
在真实环境中执行长程工具调用面临动态用户意图噪声的严重挑战。现有方法通过隐式历史扫描或文本压缩尝试提升鲁棒性,但大多假设简单场景下指令完美,不可避免地,在波动上下文下,过时约束会稀释模型注意力,引发灾难性意图偏差和无限API循环。为解决该问题,我们提出IACM-RL,一个用于鲁棒工具调用的综合框架。首先,我们引入DynamicIntent流水线,合成13种细粒度波动场景下的轨迹,并搭配一套五维诊断指标套件。其次,IACM-RL部署基于BeliefState的自生成上下文管理器,该管理器主动跟踪变化的目标,并使用结构过时标识隔离被覆盖的参数。为自主内化这种状态跟踪能力,我们采用分层意图驱动的奖励函数,结合三个辅助损失(动作校准、上下文管理器提取和状态蒸馏)来优化策略。在DynamicIntent、BFCL-V3和τ²-Bench上的实验表明,IACM-RL显著优于基线方法,减少了无限循环和过时上下文错误,同时提升了域外泛化能力。
英文摘要
Executing long-horizon tool invocations in real-world environments is severely challenged by dynamic user intent noise. Existing methods attempt robustness via implicit history scanning or text compression, yet predominantly assume perfect instructions in simplistic scenarios. Inevitably, under fluctuating contexts, obsolete constraints dilute model attention, triggering catastrophic intent deviation and infinite API loops. To resolve this, we propose IACM-RL, a comprehensive framework for robust tool invocation. First, we introduce the DynamicIntent pipeline, synthesizing trajectories across 13 fine-grained fluctuation scenarios, paired with a five-dimensional diagnostic metric suite. Second, IACM-RL deploys a BeliefState-based Self-Generated Context Manager that proactively tracks shifting goals and isolates overwritten parameters using structural stale flags. To autonomously internalize this state-tracking capability, we optimize the policy using a hierarchical intent-driven reward alongside three auxiliary losses (action calibration, CM extraction, and state distillation). Experiments on DynamicIntent, BFCL-V3, and $\mathrmτ^2$-Bench demonstrate that IACM-RL significantly outperforms baselines, reducing infinite loops and stale context errors while enhancing out-of-domain generalization.