arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.13951cs.AI

HELIX:用于递归自我改进的模型-控制器协同进化

HELIX: Model-Harness Co-evolution for Recursive Self-Improvement

Tianyu Fan, Chao Huang

首次发表
浏览论文内容

中文总结 AI 辅助

HELIX是支持模型-控制器协同进化的可溯源框架,在代码修复任务中,其65个候选组合提升了4.0%任务覆盖率,生成438条学习记录,形成模型、控制器与数据的反馈系统。

中文摘要 AI 辅助

智能体能力的扩展主要聚焦于模型的改进,但交互式智能体通过运行时控制器(harness)行动,该控制器负责调控上下文、工具、控制流和终止条件。控制器既决定了模型能够完成的任务,也影响了其学习所用的轨迹。这种耦合关系催生了用于递归自我改进的模型-控制器协同进化思路:为固定模型构建控制器,从已验证的兄弟轨迹中更新模型,并随模型能力变化重构控制器。实现该循环需要一种可控的控制器进化方式,同时保留干预的身份和效果。我们提出了HELIX,一种支持控制器进化的可溯源底层框架。HELIX将智能体系统分解为类型化端口、可复用原子、配方、产品壳和运行时策略,它使干预过程明确且可审计,同时保留轨迹、测试结果和来源。因此,控制器进化兼具两个关联作用:改进固定模型的执行效果,并生成匹配的成功案例、退化案例、近似失误和替代方案,作为后续模型改进的数据。我们在代码修复任务的一轮进化中评估了HELIX:一个包含65个候选的组合体发现了一个固定控制器,其任务覆盖率比Pi提升了4.0%;而完整组合体通过互补的兄弟行为,使已验证覆盖率最多提升了58.0%。对选定候选进行重复运行和SWE-bench评估后,一个包含200个槽位的兄弟切片生成了438条已验证的SFT、批评者、过滤和偏好记录。这些结果表明,控制器、模型和数据构成了一个反馈系统:控制器进化扩展了当前能力,并为下一轮模型改进创建学习信号;模型更新则推动下一轮控制器进化。HELIX为研究这一递归过程提供了可审计的接口,代码可在指定URL获取。

英文摘要

Scaling agent capability has largely focused on improving the model, yet an interactive agent acts through a runtime harness that mediates context, tools, control flow, and stopping. The harness shapes both what a model can accomplish and the trajectories from which it learns. This coupling motivates model-harness co-evolution for recursive self-improvement: build harnesses for a fixed model, update the model from verified sibling trajectories, and rebuild the harnesses as model capabilities change. Realizing this loop requires a controlled way to evolve harnesses while preserving intervention identity and effect. We present HELIX, a source-traceable substrate for harness evolution. HELIX decomposes agent systems into typed ports, reusable atoms, recipes, product shells, and runtime policies. It makes interventions explicit and auditable while retaining trajectories, test outcomes, and provenance. Harness evolution thus serves two linked roles: improving fixed-model execution and producing matched successes, regressions, near misses, and alternative solutions as data for subsequent model improvement. We evaluate HELIX in one evolution round on code repair. A 65-candidate portfolio discovers a fixed harness that improves task coverage by 4.0% over Pi, while the full portfolio exposes up to 58.0% more verified coverage through complementary sibling behavior. Selected candidates are assessed with repeated runs and the SWE-bench evaluator. A 200-slot sibling slice yields 438 verified SFT, critic, filter, and preference records. These results show how harness, model, and data form a feedback system: harness evolution expands current capability and creates learning signal for the next model; model updates motivate the next round of harness evolution. HELIX provides an auditable interface for studying this recursive process. Code is available at https://github.com/HKUDS/HELIX.

↑