arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

认证的长时程代码智能体进化:基于验证门控的技能优化

Certified Long-Horizon Code Agent Evolution via Validation-Gated Skill Optimization

Yifan Wang, Hao Cheng, Xiaomin Li, Yuexing Hao, Hemanth Neelgund Ramesh, Dongwon Jung, Hao Tang, Keru Wang, Chenliang Zhou, Qianhui Wu, Wenlin Yao, Ananth Grama, Andrzej Banburski-Fahey, Baolin Peng, Jaron Lanier, Jianfeng Gao

arXiv 2609.32990首次发表:更新:

AI 中文总结

本文提出VALVE验证门控框架,实现长时程代码智能体无权重更新的自我进化,在超千个SWE任务上稳定提升性能,显著降低回撤并压缩技能库。

AI 中文摘要

在无需更新模型权重的情况下实现长时程智能体自我进化,对于使已部署的智能体能够积累可复用技能并随时间改进至关重要。先前的自我进化工作主要聚焦于短时程任务,而仓库级软件工程尽管是长时程适应的理想测试平台,却仍未得到探索。在此设定下,智能体需要解决一系列连续任务,导航具有不断演化的仓库的复杂依赖关系,并持久地存储和复用经验。基于文本的技能优化为此类适应提供了一种高效的非参数化方法。然而,现有方法在长期部署中常常遭受不稳定的更新、性能回撤和智能体崩溃。在本文中,我们形式化了上下文自我进化的概念,并引入了VALVE,一个用于长时程技能优化的验证门控框架。我们确立了有限收敛性,为未来任务增益和回撤提供了理论保证,并推导了在给定容差下所需的验证和评估留出规模,具有领先阶的缩放。实验上,我们的流程VALVE在跨越超过1,000个SWE任务的进化时间跨度上实现了稳定的自我改进,在三个前沿模型(GPT-5.5、Claude-4.6和MiniMax-M2.7)上平均最终和峰值增益分别为14.9和16.5个点。验证门将平均回撤减少了75%,并产生了比无门控进化紧凑11倍的技能库。我们进一步提供了广泛的消融实验,识别了对长时程技能进化最关键的設計选择。

英文摘要

Long horizon agent self-evolution without model weight updates is essential for enabling deployed agents to accumulate reusable skills and improve over time. Prior self-evolution work has focused primarily on short-horizon tasks, while repository-level software engineering remains unexplored despite being an ideal testbed for long-horizon adaptation. In this setting, agents are required to solve streams of sequential tasks, navigate complex dependencies with evolving repositories and persistently store and reuse experience. Text-based skill optimization offers an efficient, non-parametric approach for such adaptation. However, existing methods often suffer from unstable updates, performance drawdown, and agent collapse over extended deployments. In this paper, we formalize the concept of in-context self-evolution and introduce VALVE, a validated-gated framework for long-horizon skill optimization. We establish finite convergence, provide theoretical guarantees for future-task gain and drawdown, and derive the validation and evaluation holdout sizes required for a prescribed tolerance, with leading-order scaling Empirically, our pipeline, VALVE achieves stable self-improvement over evolution horizon spanning more than 1,000 SWE tasks, with average final and peak gains of $14.9$ and $16.5$ points across three frontier models (GPT-5.5, Claude-4.6 and MiniMax-M2.7). The validation gate reduces average drawdown by 75% and produces an 11x more compact skill bank than ungated evolution. We further present extensive ablations identifying the design choices most critical to long-horizon skill evolution.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑