AI 中文总结
本研究推出PAST-Bench基准评测个人智能体递归自我改进能力,发现改进不均,在此基础上开发Hermes+提升收益,为相关研究提供评估诊断基础
AI 中文摘要
递归自我改进要求智能体将积累的经验转化为未来更优的行为。个人AI智能体为研究该能力提供了具体场景,因为它们会在会话间保留偏好、任务历史、工具例程和学习到的技能。然而,保留的经验是否真的能让智能体随时间推移得到改进,尚未得到系统测试。我们推出PAST-Bench,这是一个旨在隔离该问题的基准。每个智能体在匹配条件下运行一系列新会话任务,通过开启和关闭保留的经验来控制变量。该基准涵盖记忆、程序复用、信息收集和更新等方面的26种场景和204个 episodes(任务片段)。我们同时报告后续任务的收益,以及这些收益是否遵循预期的保存、检索和更新路径。在7种基础模型和4种智能体框架中,改进是真实存在的,但在不同能力上表现不均。具有相同总体收益的智能体,其收益是否由预期路径的证据支持可能存在显著差异。基于这些发现,我们开发了Hermes+,它在智能体循环的各个阶段对Hermes进行了五项针对性干预。Hermes+提高了保留经验带来的平均收益,并提供了更清晰的路径证据,在需要替换过时状态的任务上改进最为显著,不过该效果仍取决于能力和模型。总体而言,PAST-Bench和Hermes+为研究持久智能体如何从保留经验发展为通过经验实现系统性改进提供了评估和诊断基础。代码:this https URL
英文摘要
Recursive self-improvement requires agents to turn accumulated experience into better future behavior. Personal AI agents offer a concrete setting for studying this capability because they retain preferences, task histories, tool routines, and learned skills across sessions. Yet whether retained experience actually improves them over time has not been systematically tested. We introduce PAST-Bench, a benchmark designed to isolate this question. Each agent runs through ordered sequences of fresh-session tasks under matched conditions that turn retained experience on and off. It spans 26 scenarios and 204 episodes across memory, procedural reuse, information gathering, and update. We report both later-task gains and whether those gains follow the intended save, retrieve, and update pathway. Across seven base models and four agent frameworks, improvement is real but uneven across capabilities. Agents with the same headline gain can differ markedly in whether that gain is supported by evidence of the intended pathway. Guided by these findings, we develop Hermes+, which extends Hermes with five targeted interventions across stages of the agent loop. Hermes+ raises the average gain from retained experience and provides clearer pathway evidence, with its strongest improvement on tasks requiring outdated state to be replaced, although the effect remains capability- and model-dependent. Together, PAST-Bench and Hermes+ provide an evaluation and diagnostic foundation for studying how persistent agents can progress from retaining experience to systematically improving through it. Code: https://github.com/Gen-Verse/PAST-Bench
CommentsCode: https://github.com/Gen-Verse/PAST-Bench