arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10178cs.SE

一个方案,多种适配:自演化在不同语言和模型中编码了什么

One Recipe, Many Harnesses: What Self-Evolution Encodes Across Languages and Models

Siqi Yang, Qianlan Yang, Yu-Xiong Wang, Saurabh Pujar, Martin Hirzel

首次发表
浏览论文内容

中文总结 AI 辅助

该研究固定演化方案,在八种编程语言和三个基础模型上分析自演化适配工具,发现其为语言需求与模型差距共同塑造的补偿层,可提升编码智能体性能并能蒸馏为通用工具。

中文摘要 AI 辅助

自演化适配工具(harness)是一种闭环系统,其中智能体会检查自身的执行轨迹并编辑其提示、工具和记忆。这类工具能可靠地在评估中提升编码智能体的性能,但现有研究仅报告了总体收益,未分析演化后的工件编码了什么。因此仍不清楚它们是编码了针对特定基准的适配、特定语言的工程知识,还是对底层模型局限性的补偿。我们通过在八种编程语言(Multi-SWE-Bench)和三个基础模型组成的网格中固定演化方案,并分析生成的适配工具,来拆解这些因素。该方案通过类型化的失败信号处理每一次编辑,并将其记录为可证伪的契约,使得每次修改都可在演化后追溯。研究得出四个发现:(1) 该循环在大多数情况下,相较于最小种子和手动设计的mini-SWE-agent框架,提升了保留的解决率,但存在两个无收益区域;(2) 收益补偿可恢复的执行缺陷,当缺陷质量接近零时收益也接近零,且主导缺陷因具体情况而异,适配工具缩小了策略能做与实际做的差距;(3) 演化后的适配工具在不同语言间共享抽象操作手册,但用几乎不相交的语言生态系统机制实现;(4) 共享核心可迁移并能被蒸馏为一个通用适配工具,而生态系统边缘则对此有抗性,需要原生重新演化。这些结果共同将演化后的适配工具重新定义为一种清晰的补偿层,由语言的工程需求和模型的行为差距共同塑造,而非不透明的基准调优框架。

英文摘要

Self-evolving harnesses are closed-loop systems in which an agent inspects its own rollouts and edits its prompts, tools, and memory. They reliably improve coding agents in evaluations, but prior work reports aggregate gains rather than analyzing what the evolved artifacts encode. It therefore remains unclear whether they encode benchmark-specific adaptations, language-specific engineering knowledge, or compensation for limitations of the underlying model. We disentangle these factors by holding an evolution recipe fixed across a grid of eight programming languages (Multi-SWE-Bench) and three base models, and analyzing the resulting harnesses. The recipe routes every edit through a typed failure signal and records it as a falsifiable contract, making each modification attributable after evolution. Four findings emerge. (1)The loop improves held-out solve rates over both a minimal seed and the manually designed mini-SWE-agent scaffold in most cells, but with two null regions. (2)Gains compensate recoverable execution defects, where defect mass is near zero, and gain is near zero; which defect dominates is cell-specific. A harness closes the gap between what a policy can do and what it does. (3)Evolved harnesses share an abstract playbook across languages but instantiate it with almost disjoint language ecosystem machinery. (4)The shared core transfers and can be distilled into one universal harness, while an ecosystem margin resists both and requires native re-evolution. Together, these results recast the evolved harness as a legible compensation layer, shaped jointly by the language's engineering demands and the model's behavioral gaps, rather than an opaque benchmark-tuned scaffold.

补充信息

↑