arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.09629cs.AI

反思自进化智能体:我们仍需要规定优化流程吗?

Rethinking Self-Evolving Agents: Do We Still Need Prescribed Optimization Pipelines?

Hui Xue, Fan Yang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出开放端优化(OEO),对比SkillOpt、GEPA等规定优化流程,发现GPT-5.5驱动的OEO在多数场景优于规定流程,且资源消耗更低,表明足够强大的优化器可自主构建改进路径。

中文摘要 AI 辅助

自进化智能体通常围绕规定优化流程构建:该框架决定如何收集证据、修改持久化构件、选择候选方案以及停止优化。我们提出,当前沿模型充当优化器时,这种特定任务的流程是否仍有必要。我们引入开放端优化(Open-Ended Optimization, OEO),其在保持目标、允许的交互、资源预算、数据边界和评估固定的同时,允许优化器在线组合改进过程。我们将OEO与两种互补的规定方法对比:SkillOpt(具有有限编辑的分阶段流程)和GEPA(一种反思式进化搜索)。在8个基准目标模型设置的14次头对头比较中,GPT-5.5驱动的OEO取得12次胜利、1次平局,以及1次仅差0.21个百分点的微弱失败。它使用的中位数仅为SkillOpt配置的目标交互token预算的34.3%。一次单轮、零交互的对照实验表明,该增益并非由单一的先验驱动重写所解释。然而,委托存在能力边界:当优化器能力中等时,SkillOpt的表现优于OEO;而能力较弱的优化器无法通过未改变的OEO接口运行。在完全装备的OEO-SkillOpt配对中,轨迹分析进一步显示,规定对优化流程的改变比对最终行为的改变更具一致性。这些发现共同将规定流程重塑为依赖能力的支撑结构:必要的约束仍为外部约束,但足够强大的优化器可从可测量反馈出发,自主构建通向持久改进的路径。

英文摘要

Self-evolving agents are usually built around prescribed optimization pipelines: the framework decides how to gather evidence, revise a persistent artifact, select candidates, and stop. We ask whether this task-specific procedure remains necessary when a frontier model acts as the optimizer. We introduce Open-Ended Optimization (OEO), which keeps the objective, permitted interactions, resource budget, data boundary, and evaluation fixed while allowing the optimizer to compose the improvement process online. We compare OEO with two complementary prescribed approaches: SkillOpt, a staged pipeline with bounded edits, and GEPA, a reflective evolutionary search. Across 14 head-to-head comparisons over 8 benchmark-target-model settings, GPT-5.5-driven OEO records 12 wins, 1 tie, and 1 narrow loss of 0.21 percentage points. It uses a median 34.3 percent of SkillOpt's configured target-interaction token budget. A one-shot, zero-interaction control shows that the gains are not explained by a single prior-driven rewrite. However, delegation has a capability boundary: SkillOpt outperforms OEO with a medium optimizer, and a weak optimizer cannot operate through the unchanged OEO interface. In the fully instrumented OEO-SkillOpt pair, trajectory analysis further shows that prescription changes how optimization proceeds more consistently than it changes final behavior. Together, these findings recast prescribed pipelines as capability-dependent scaffolding: essential constraints remain external, but a sufficiently capable optimizer can compose the route from measurable feedback to persistent improvement.

发表机构

  • Microsoft Research(微软研究院)

机构由 AI 辅助整理,请以论文原文为准。

↑