arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向长 horizon 智能体利用的递归经验-工作记忆演化

Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses

Zhaochen Yu, Yingcheng Wu, Zhenfei Yin, Kaiyuan Chen, Zhe Zhao, Mengdi Wang, Shuicheng Yan, Ling Yang

arXiv 2608.24876首次发表:更新:

发表机构

Princeton University; Stanford University; University of Oxford(普林斯顿大学; 斯坦福大学; 牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对长 horizon 智能体的递归自我改进难题,提出 Recuris 架构,通过递归记忆演化提升任务成功率,在多个基准和模型上实现显著性能提升,为 RSI 提供可扩展基础。

AI 中文摘要

递归自我改进(RSI)在长 horizon 任务中仍然极具挑战性,不断增长的历史会模糊任务状态并导致技能调用错位。我们提出 Recuris,一种面向长 horizon 智能体利用的递归经验-工作记忆架构,其中工作记忆跟踪任务进度并从经验记忆中引导技能选择,使技能使用基于当前需求而非完整历史。这种耦合还将执行转化为结构化证据,可将故障定位到特定记忆组件。在所有任务中,一个固定的元智能体将该证据转化为对技能记忆的局部化、经验证门控的更新,这些更新会重塑执行并产生新证据,形成一个有界的递归记忆演化循环。在四个长 horizon 基准和十个模型上,Recuris 在 37 个完成的模型-基准对中,有 35 个提升了任务成功率,使前沿模型达到了 SOTA 级别的任务成功率:在 tau-bench 上,它为 GPT-5.6 Sol 增加了 17.8 个百分点,为 Claude Opus 5 增加了 15.6 个百分点,使 Opus 5 的任务成功率达到 87.9%;在 SkillFlow 上,为 Qwen3.6-27B 和 Qwen3.6-35B 分别增加了 16.6 和 13.5 个百分点。随着交互 horizon 的增长,优势进一步扩大,在最长任务上达到 32.2 个百分点,常见的长 horizon 故障最多降低了 80%。这些结果表明,递归演化的记忆可作为 RSI 的可扩展基础,使智能体能够将积累的经验持续转化为日益有效的长 horizon 行为。代码:this https URL

英文摘要

Recursive self-improvement (RSI) remains hard in long-horizon tasks, where growing histories obscure the task state and misalign skill invocation. We introduce Recuris, a recursive Experiential-Working Memory architecture for long-horizon agent harnesses, in which Working Memory tracks task progress and guides skill selection from Experiential Memory, grounding skill use in current needs rather than the full history. This coupling also turns execution into structured evidence that localizes failures to specific memory components. Across tasks, a fixed Meta-Agent turns that evidence into localized, validation-gated updates to Skill Memory that reshape execution and yield new evidence, forming a bounded recursive memory-evolution loop. Across four long-horizon benchmarks and ten models, Recuris improves task success in 35 of the 37 completed model-benchmark pairs, carrying frontier models to SOTA-level task success: on tau-bench it adds +17.8 points to GPT-5.6 Sol and +15.6 to Claude Opus 5, taking Opus 5 to 87.9%, and +16.6/+13.5 points on Qwen3.6-27B/35B on SkillFlow. The advantage widens as the interaction horizon grows, to +32.2 points on the longest tasks, and common long-horizon failures fall by up to 80%. These results position recursively evolving memory as a scalable foundation for RSI, enabling agents to continuously transform accumulated experience into increasingly effective long-horizon behavior. Code: https://github.com/Gen-Verse/Recuris

CommentsCode: https://github.com/Gen-Verse/Recuris

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑