发表机构
Zhejiang University; Alibaba Group(浙江大学; 阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对自进化智能体仅凭端点评估不全面的问题,提出EvoPathBench基准,在过程级追踪能力演化,发现分布偏移下泛化增益减弱、保持损失集中且规则适应不可靠,并指出候选评估与选择是改进关键。
AI 中文摘要
自进化智能体将交互反馈转化为持久性产物,如记忆或技能,这些产物进而指导后续决策。随着这些产物在经验流中被迭代更新,它们所支持的能力也可能随之演化。因此,仅凭端点性能无法全面反映自进化过程。过程级评估对于识别目标能力何时出现以及后续更新是增强、保持还是削弱该能力至关重要。基于此,我们提出了EvoPathBench,一个在产物级自进化过程中追踪个体能力的基准。EvoPathBench固定基础模型和工具,在连续检查点冻结演化中的产物,并在保留的回合上评估目标能力。该基准使用公开交易数据和校准轨迹来评估智能体自进化。它测试三种能力:对未见任务的泛化、无关学习后的保持以及对新证据的规则适应。实验结果表明,在相似未见任务上的增益在分布偏移下往往减弱,保持损失集中在少数演化路径中,且没有方法实现可靠的规则适应。此外,尽管自进化使智能体能够生成具有显著保留增益的候选产物,但所选更新始终未能实现这一潜力。综上所述,这些发现确立了能力级过程评估作为分析自进化的基础,并识别出候选评估与选择作为改进的关键目标。
英文摘要
Self-evolving agents convert interaction feedback into persistent artifacts, such as memories or skills, which in turn guide subsequent decisions. As these artifacts are iteratively updated throughout an experience stream, the capabilities they support may evolve. Consequently, endpoint performance alone offers an incomplete view of self-evolution. Process-level evaluation is therefore essential to identify when a target capability emerges and whether later updates strengthen, preserve, or weaken it. Motivated by this, we propose \textsc{EvoPathBench}, a benchmark that tracks individual capabilities during artifact-level self-evolution. EvoPathBench fixes the base model, tools, freezes evolving artifacts at successive checkpoints, and evaluates the target capability on held-out episodes. This benchmark evaluates agent self-evolution using public trading data and calibrated trajectories. It tests three capabilities: generalization to unseen tasks, retention after unrelated learning, and rule adaptation to new evidence. Experimental results show that gains on similar unseen tasks often weaken under distribution shift, retention losses are concentrated in a minority of evolution paths, and no method achieves reliable rule adaptation. Moreover, while self-evolution enables agents to generate candidate artifacts with substantial held-out gains, the selected updates consistently fall short of realizing this potential. Together, these findings establish capability-level process evaluation as a foundation for analyzing self-evolution, identifying candidate evaluation and selection as key targets for improvement.