arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EVOHARNESSBENCH:你的智能体能跟上不断演进的工具链吗?

EvoHarnessBench: Can Your Agents Keep Pace with an Evolving Harness?

Zixuan Ke, Vaidehi Patil, Haizhou Shi, Yang Li, Ye Liu, Sarath Shekkizhar, Anurag Koul, Jiayu Wang, Xuan Phi Nguyen, Semih Yavuz, Mohit Bansal, Shafiq Joty

arXiv 2609.04280首次发表:更新:

发表机构

Salesforce Research; University of North Carolina at Chapel Hill; University of Wisconsin–Madison(赛富时研究院; 北卡罗来纳大学教堂山分校; 威斯康星大学麦迪逊分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出EVOHARNESSBENCH基准,评估智能体在工具链沿工具、技能、智能体三维演进时的表现,发现工具链扩展会引发遗忘、自演进适应收益不一致等问题,为相关智能体构建提供了新挑战。

AI 中文摘要

基于大语言模型(LLM)的现代智能体通过工具、可复用技能及专用智能体构成的工具链运行,该工具链决定了智能体的观测内容与可执行操作。实际应用中,随着新能力的加入,该工具链会持续演进。本文提出了EVOHARNESSBENCH,这是一个用于在工具链沿工具、技能、智能体三个维度受控演进的场景下评估智能体的基准。与现有的智能体持续学习基准不同,现有基准通常将非平稳性(即随时间变化的内容)置于任务流中而保持工具链固定,EVOHARNESSBENCH则将非平稳性置于外部提供的工具链本身。该基准包含17个多阶段工具链流,由基于验证器的基准确定性构建而成,涵盖802个任务、520个工具、42项技能及62个智能体。我们针对工具链演进的核心挑战设置了两种互补的评估场景:部署评估,用于隔离工具链扩展时对先前可及能力的保留情况;自演进适应评估,用于测试引入新能力时积累的经验是否仍有用。我们的结果揭示了三个持续存在的差距:第一,仅工具链扩展就会降低先前已解决任务的性能,产生工具链诱导的遗忘;第二,自演进适应的收益在工具链演进阶段、能力维度及环境间仍不一致;第三,保留与适应可能存在方向冲突:保留早期能力不一定能提升对新引入能力的适应,反之亦然。这些结果表明,工具链演进是构建能跟上不断演进的工具链同时保留先前有效行为的智能体的一个独特挑战。

英文摘要

Modern LLM-based agents operate through a harness of tools, reusable skills, and specialist agents that shapes what they observe and what they can do. In practice, this harness continually evolves as new capabilities are added. We introduce EvoHarnessBench, a benchmark for evaluating agents under controlled harness evolution across three axes (tools, skills, and agents). Unlike existing continual-learning benchmarks for agents, which typically place non-stationarity (i.e., what changes over time) in the task stream while keeping the harness fixed, EvoHarnessBench places non-stationarity in the externally supplied harness itself. It contains 17 multi-stage harness streams constructed deterministically from verifier-based benchmarks, comprising 802 tasks, 520 tools, 42 skills, and 62 agents. We evaluate two complementary settings corresponding to the central challenges of harness evolution: deployment evaluation, which isolates retention of previously accessible competence as the harness expands, and self-evolving adaptation evaluation, which tests whether accumulated experience remains useful as new capabilities are introduced. Our results reveal three persistent gaps. First, harness expansion alone can degrade performance on previously solved tasks, producing harness-induced forgetting. Second, gains from self-evolving adaptation remain inconsistent across stages of harness evolution, capability axes, and environments. Third, retention and adaptation can pull in different directions: preserving earlier competence does not necessarily improve adaptation to newly introduced capabilities, and vice versa. These results establish harness evolution as a distinct challenge for building agents that can keep pace with an evolving harness while preserving previously effective behavior.

Commentshttps://mas-orchestra.salesforceresearch.ai/evoharness/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑