arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自进化智能体:基于似然引导的工具空间优化

Self-Evolving Agents via Likelihood-Guided Tool-Space Optimization

Xuanqi Zhang, Ruinan Jin, Running Yang, Yuxuan Zhang, Minghui Chen, Wenlong Deng, Xiaoxiao Li

arXiv 2609.34151首次发表:更新:

发表机构

University of British Columbia; Vector Institute; Columbia University(不列颠哥伦比亚大学; 向量研究所; 哥伦比亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对工具空间自进化中输出不可见、跨请求无状态及推理成本高的问题,提出基于似然贡献评分的LOTS方法,在保持模型参数不变下从累积经验优化工具空间,提升性能并减少上下文。

AI 中文摘要

自进化智能体能够持续改进自身行为,而工具则定义了它们与环境交互的可执行动作空间。然而,将完整的工具库暴露给模型会引入大量无关上下文,并可能损害工具使用决策。我们研究工具空间的自进化,其中每个重复出现的任务类型维护一个由累积输出经验构建的持久工具空间。我们识别出现有方法的三个局限性:(1)不考虑输出的选择:它们主要依赖工具描述或模型先验,而非观察到的工具输出;(2)跨请求无状态:它们为每个请求独立选择工具,未将先前的输出经验整合为持久的任务特定状态;(3)推理成本:对于同一任务的后续请求,它们反复搜索、排序或推理候选工具。我们通过输出感知的工具评分、持久的任务特定工具空间、摊销的工具选择以及跨模型可复用的配置来解决这些局限性。我们引入LOTS(仅似然工具评分),它从累积输出经验中进化智能体的工具空间,同时保持模型参数固定。在每个请求后,LOTS保留模型生成的答案,并通过测量当移除其观察到的输出时答案似然的变化来估计每个工具的贡献。这些贡献在每个重复任务内聚合,以对工具进行排序并更新其持久空间。在三个基准上,LOTS提升了任务性能,同时大幅减少了工具上下文。更重要的是,序列实验表明任务特定空间持续存在并随时间不断改进,而跨模型实验表明学习到的配置可迁移到不同模型。

英文摘要

Self-evolving agents can continually improve their behavior, while tools define the executable action space through which they interact with the environment. However, exposing the full tool library to model introduces substantial irrelevant context and can impair tool-use decisions. We study tool-space self-evolution, where each recurring task type maintains a persistent tool space which is constructed from accumulated output experience. We identify three limitations of existing methods: (1) output-unaware selection: they rely primarily on tool descriptions or model priors rather than observed tool outputs; (2) statelessness across request: they select tools independently for each request without consolidating prior output experience into persistent task-specific state; (3) inference cost: they repeatedly search, rank, or reason over candidate tools for subsequent requests of the same task. We address these limitations through output-aware tool scoring, persistent task-specific tool spaces, amortized tool selection, and reusable configurations across models. We introduce LOTS (Likelihood-Only Tool Scoring), which evolves an agent's tool space from accumulated output experience while keeping model parameters fixed. After each request, LOTS holds the model's generated answer and estimates each tool's contribution by measuring how much the answer likelihood changes when its observed output is removed. These contributions are aggregated within each recurring task to rank tools and update its persistent space. Across three benchmarks, LOTS improves task performance while substantially reducing tool context. More importantly, sequential experiments demonstrate that task-specific spaces persist and continue to improve over time, while cross-model experiments show that learned configurations transfer across different models.

Comments39 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑