发表机构
University of Science and Technology of China(中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对工具智能体修复中重复操作选择导致的高计算成本问题,提出ReCommit框架,利用掩码扩散语言模型的单次并行读出评分引导分层搜索,在Agent-Diff基准上显著提升恢复率并降低修复时间。
AI 中文摘要
工具智能体通过大型语言模型借助外部工具执行操作,然而,即使工具调用执行成功,用户请求仍可能未得到满足。工具智能体修复旨在寻找能够成功执行并满足原始请求的替代调用序列。然而,修复需要同时探索操作选择及其具体实现,这使得完整序列的重新生成代价高昂。此外,当失败源于操作的实现方式时,重新生成会重复进行操作选择。由此带来的挑战是,在保留对替代操作及其实现探索的同时,减少这种重复。因此,我们将修复形式化为对操作支持集的分层搜索,操作支持集我们定义为允许的操作类型集合,这些集合定义了具体工具调用序列的可复用搜索区域。我们提出了ReCommit,一个无需训练、扩散引导的框架,用于在减少修复计算量的同时改进工具智能体的失败恢复。ReCommit通过复用来自掩码扩散语言模型单次并行读出的操作类型评分,将操作级别的提议计算分摊到多次修复尝试中。这些评分引导跨支持集的搜索,而实现搜索则在每个支持集内探索替代的实体绑定、参数和动作组合。在Agent-Diff基准中四个企业服务的真实失败上进行的实验表明,在修复预算B=3和B=13时,相对于评估过的最强8B对比方法,恢复率分别相对提升75.9%和63.2%,平均全预算修复时间分别减少61.3%和51.3%。ReCommit实现了有利的恢复-成本权衡,包括在与评估过的32B模型的比较中。
英文摘要
Tool agents use large language models to act through external tools, yet successfully executed calls can still leave user requests unfulfilled. Tool-agent repair seeks alternative call sequences that execute successfully and fulfill the original requests. However, repair requires exploring both operation choices and their concrete realizations, making complete-sequence regeneration costly. Moreover, regeneration repeats operation selection even when failure arises from how those operations are realized. The resulting challenge is to reduce this repetition while preserving exploration of alternative operations and realizations. Therefore, we formulate repair as hierarchical search over operation supports, which we introduce as sets of permitted operation types that define reusable search regions for concrete tool-call sequences. We propose ReCommit, a training-free, diffusion-guided framework for improving tool-agent failure recovery while reducing repair computation. ReCommit amortizes operation-level proposal computation across repair trials by reusing operation-type scores from a single parallel readout of a masked diffusion language model. These scores guide search across supports, while realization search explores alternative entity bindings, arguments, and action composition within each support. Experiments on real failures across four enterprise services in the Agent-Diff benchmark show 75.9\% and 63.2\% relative recovery gains with 61.3\% and 51.3\% reductions in mean full-budget repair time at repair budgets $B=3$ and $B=13$, respectively, over the strongest evaluated 8B comparison method. ReCommit achieves a favorable recovery--cost trade-off, including in comparisons with the evaluated 32B models.