PATH-Bench:面向终身智能体的路径依赖评估基准
PATH-Bench: Path-Dependent Evaluation of Lifelong Agents
浏览论文内容
中文总结 AI 辅助
本研究构建了面向终身智能体路径依赖评估的基准PATH-Bench,评估8个智能体后提出SEU机制,该机制可减少遗忘、提升正向迁移,为设计更优终身智能体提供指导。
中文摘要 AI 辅助
当前,基于大语言模型(LLM)的终身智能体日益通过外部学习状态实现适配,这类状态会将过往交互存储为可检索记忆或可复用技能,但现有基准极少考量累积经验的路径如何影响智能体的知识迁移与保留。本研究构建了PATH-Bench——一个用于终身智能体路径依赖评估的基准,该基准通过多模型上下文学习估计定向任务关系,构建以探测任务为核心且包含可控有益与干扰经验的序列,并重复评估探测任务以测量平均性能、正向迁移、反向迁移及遗忘情况。我们在正向主导与负向主导经验下,针对单轮代码生成与多轮工具使用任务,评估了8个代表性智能体。基准结果显示,经验效用同时取决于经验的表征方式与任务的交互结构,强迁移能力不代表具备良好的经验保留能力,且后期经验可重塑学习路径中前期获得的收益。基于上述发现,我们提出了选择性经验利用框架(Selective Experience Use, SEU),该智能体机制可调控路径累积的经验对每个新任务的影响,保留有益经验项并过滤潜在干扰项。在多数场景中,SEU可持续减少遗忘并提升正向迁移效果。PATH-Bench既提供了可控的评估框架,也为设计更具选择性与鲁棒性的终身智能体提供了可操作的指导。
英文摘要
Lifelong LLM agents increasingly adapt through external learning states that store past interactions as retrievable memories or reusable skills, yet existing benchmarks rarely account for how the path of accumulated experience shapes what agents transfer and retain. In this work, we establish PATH-Bench, a benchmark for path-dependent evaluation of lifelong agents. PATH-Bench estimates directed task relationships via multi-model in-context learning, constructs probe-centered sequences with controlled helpful and interfering histories, and repeatedly evaluates probe tasks to measure average performance, forward transfer, backward transfer, and forgetting. We evaluate eight representative agents on single-turn code generation and multi-turn tool-use tasks under positive- and negative-dominant histories. Benchmark results show that experience utility depends jointly on how experience is represented and on the task's interaction structure, that strong transfer does not ensure retention, and that later experience can reshape gains acquired earlier in the learning path. Based on these findings, we propose Selective Experience Use (SEU), an agent harness that regulates how path-accumulated experience influences each new task, admitting helpful items while filtering out potential interference. SEU consistently reduces forgetting while improving forward transfer in the majority of settings. The PATH-Bench provides both a controlled evaluation framework and actionable guidance for designing more selective and robust lifelong agents.