arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.09380cs.AI

OpenLoopEvolve:面向长周期复杂任务中循环策略的可验证自进化框架

OpenLoopEvolve: A Verifiable Self-Evolution Framework for Loop Policies in Long-Horizon Complex Tasks

Siqi Wang, Xinlin Li, Zhenglin Li, Li Li

首次发表
浏览论文内容

中文总结 AI 辅助

OpenLoopEvolve是面向长周期复杂任务的可验证自进化框架,通过在线与离线两种进化模式优化循环策略,在YC-Bench基准上提升了任务性能、成功率与风险指标。

中文摘要 AI 辅助

长周期复杂任务要求智能体在持续变化的环境中反复观测状态、制定计划、调用工具、验证结果并从失败中恢复。然而这类控制经验往往局限于单一上下文或固定提示词,难以跨历史轨迹积累和复用。本文提出OpenLoopEvolve(OLE),一种以循环策略为核心的自进化框架。OLE将智能体的观测、规划、记忆、动作、验证、恢复、终止及预算控制表示为带版本和谱系的可移植策略资产,并提供可根据实际需求选择的在线与离线进化模式:在线模式基于持续运行的反馈触发候选生成,离线模式从归档轨迹和失败证据中搜索候选策略。两种模式共享一套进化机制,包括大语言模型自主提案、“优胜者-挑战者”配对评估及稳健发布。在线发布的策略会在后续任务边界激活,通过后续反馈监控,当性能下降时回滚至父版本。在模拟业务基准YC-Bench上,两种模式相较于固定初始循环策略均提升了聚合任务性能、任务成功率及风险指标。结果表明,将循环策略视为可管控资产可支持控制经验的积累、比较、发布与复用,提升智能体在长周期复杂任务上的性能。

英文摘要

Long-horizon complex tasks require agents to repeatedly observe states, formulate plans, invoke tools, verify results, and recover from failures in continuously changing environments. However, such control experience often remains confined to a single context or a fixed prompt, and is difficult to accumulate and reuse across historical traces. This paper presents OpenLoopEvolve (OLE), a self-evolution framework centered on the Loop Policy. OLE represents an agent's observation, planning, memory, action, verification, recovery, stopping, and budget control as portable policy assets with versions and lineages, and provides online and offline evolution modes that can be selected according to practical needs: the online mode triggers candidate generation based on feedback from continuous operation, whereas the offline mode searches for candidate policies from archived traces and failure evidence. Both modes share an evolution mechanism consisting of autonomous proposals by a large language model, Champion--Challenger paired evaluation, and robust release. Policies released online are activated at a subsequent task boundary, monitored using subsequent feedback, and rolled back to their parent versions when degradation conditions are met. On the simulated business benchmark YC-Bench, both modes improve aggregate task performance, task success rate, and risk metrics relative to a fixed initial Loop Policy. The results indicate that treating the Loop Policy as a governable asset can support the accumulation, comparison, release, and reuse of control experience and improve agent performance on long-horizon complex tasks.

发表机构

  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

↑