AI 中文总结
本研究针对智能体工具使用鲁棒性不足的问题,提出ExpG机制,经实验验证其可提升多类工具任务性能,还能让小模型智能体超越未用该机制的大模型智能体。
AI 中文摘要
当前智能体的性能瓶颈正逐渐从模型能力转向执行过程的鲁棒性。工具是智能体与外部环境交互的主要接口,然而现有方法很少关注如何在多样的运行时条件下确保工具使用的鲁棒性。为解决这一问题,我们提出ExpG机制,该机制构建并优化自适应指导,捕捉每个工具的能力边界与最佳实践,从而让智能体更鲁棒、更有效地使用工具。ExpG包含三个阶段:(1)经验获取,从历史执行轨迹中分析工具调用质量,通过多维度归因生成结构化的可学习经验;(2)经验提炼,通过过滤无效经验、基于等价类方法选择代表性经验并将其总结为可泛化的指导,维持经验池的有效性;(3)经验复用,在未来任务解决过程中自适应应用该指导。大量实验表明,ExpG在工具选择、工具调用和响应生成任务中均实现了持续提升,能让较小规模的智能体超越未使用ExpG的更大规模智能体。此外,ExpG在具有挑战性的设置中取得了尤为显著的增益,为实现更鲁棒的工具使用提供了有前景的路径。我们的代码、实验和结果均已公开。
英文摘要
The performance bottleneck of agents is increasingly shifting from model capability to the robustness of their execution processes. Tools play a central role as the primary interface through which agents interact with external environments, yet existing methods rarely focus on ensuring robust tool use across diverse runtime conditions. To address this problem, we propose ExpG, a mechanism that builds and refines adaptive guidance capturing each tool's capability boundaries and best practices, thereby enabling agents to use tools more robustly and effectively. ExpG consists of three phases: (1) experience acquisition, which analyzes tool invocation quality from historical execution trajectories, producing structured learnable experiences through multi-aspect attribution; (2) experience distillation, which keeps the experience pool effective by filtering unhelpful experiences, selecting representative ones with an equivalence-class-based method, and summarizing them into generalizable guidance; and (3) experience reuse, which applies the guidance adaptively during future task solving. Extensive experiments show that ExpG brings consistent improvements across the tool selection, tool calling, and response generation tasks, enabling smaller agents to outperform larger ones that do not use ExpG. Moreover, ExpG achieves particularly strong gains in challenging settings, suggesting a promising path toward more robust tool use. Our code, experiments, and results are available.
CommentsPreprint.14 figures