arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.14857cs.CL

ModularRSI:模块化且可泛化的递归式框架自改进

ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement

Siwei Wu, Jincheng Ren, Yizhi Li, Haau-Sing Li, Chengran Yang, Yuxuan Zhang, Weicheng Gu, Jian Yang, Riza Batista-Navarro, Chuanyi Zhang, Xianglong Liu, Ming Zh… 展开作者

Siwei Wu, Jincheng Ren, Yizhi Li, Haau-Sing Li, Chengran Yang, Yuxuan Zhang, Weicheng Gu, Jian Yang, Riza Batista-Navarro, Chuanyi Zhang, Xianglong Liu, Ming Zhou, Bryan Dai, Chenghua Lin

首次发表
浏览论文内容

中文总结 AI 辅助

针对框架递归自改进难以泛化的问题,提出ModularRSI,通过对比成功与失败轨迹并模块化独立演化五个功能模块,在基准外任务上训练,实现跨任务和跨模型的稳定改进。

中文摘要 AI 辅助

近期工作将递归自改进(RSI)扩展到智能体框架(agent harnesses),用于长时程编码和终端任务,使智能体能够从经验中改进执行机制。然而,可泛化的框架RSI仍然具有挑战性。首先,在评估基准或其子集上演化框架,难以区分可复用的改进与针对基准的特定适配。其次,单轨迹更新可能将系统性的框架缺陷与实例特定的推理和解决方案细节混为一谈,产生对未见任务迁移性较差的修改。第三,在单体框架中定位重复出现的行为缺陷较为困难,而整体框架优化可能纠缠不相关的机制,使归因和验证复杂化。我们提出ModularRSI,一个与基准不相交的、对比式的、模块化的框架演化框架。ModularRSI对比同一任务的成功与失败轨迹,并跨任务聚合证据以识别重复出现的行为缺陷。它将可演化的框架分解为五个功能模块:智能体循环(Agent Loop)、工具使用(Tool Use)、观测管理(Observation Management)、上下文管理(Context Management)和任务完成检测(Task Completion Detection)。每个模块在受限的修改范围内独立演化,随后通过集成阶段将演化后的模块组合成统一框架并解决潜在冲突。为支持与基准不相交的演化,我们从外部来源整理了2,000个可执行的演化任务,这些任务与下游评估基准不相交。在TB2.0和SWE-Bench Verified上的实验表明,在未见过的领域内和跨领域任务上均取得了一致的改进,且演化后的框架还能在不同基础模型之间迁移。

英文摘要

Recent work extends recursive self-improvement (RSI) to agent harnesses for long-horizon coding and terminal tasks, enabling agents to improve execution mechanisms from experience. However, generalizable harness RSI remains challenging. First, evolving harnesses on evaluation benchmarks or their subsets makes it difficult to distinguish reusable improvements from benchmark-specific adaptation. Second, single-trajectory updates can conflate systematic harness deficiencies with instance-specific reasoning and solution details, producing modifications that transfer poorly to unseen tasks. Third, localizing recurring behavioral deficiencies within monolithic harnesses is difficult, while whole-harness optimization can entangle unrelated mechanisms and complicate attribution and validation. We propose ModularRSI, a benchmark-disjoint, contrastive, and modular framework for generalizable harness evolution. ModularRSI contrasts successful and failed trajectories for the same task and aggregates evidence across tasks to identify recurring behavioral deficiencies. It decomposes the evolvable harness into five functional modules: Agent Loop, Tool Use, Observation Management, Context Management, and Task Completion Detection. Each module evolves independently within a restricted modification scope, followed by an integration stage that combines the evolved modules into a unified harness and resolves potential conflicts. To support benchmark-disjoint evolution, we curate 2,000 executable evolution tasks from external sources that are disjoint from downstream evaluation benchmarks. Experiments on TB2.0 and SWE-Bench Verified show consistent improvements on unseen in-domain and cross-domain tasks, with the evolved harness also transferring across different foundation models.

发表机构

  • Beihang University(北京航空航天大学)
  • University of Manchester(曼彻斯特大学)
  • IQuest Research
  • M-A-P
  • Langboat(澜舟科技)
  • Hohai University(河海大学)

机构由 AI 辅助整理,请以论文原文为准。

↑