arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28363cs.AI

EvoUndo:面向大语言模型智能体执行框架的可恢复性约束自进化

EvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses

Tanmay Sah, Dolly Sah, Harshul Jain, Tanya Sah

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对LLM智能体自进化的可恢复性问题,提出EvoUndo框架,通过协同设计验证、状态落地等,在600个任务中实现高比例的失败变更恢复,且效应存在模型依赖性。

中文摘要 AI 辅助

大语言模型(LLM)智能体越来越多地在运行时修改自身的提示词、工具、中间件、资源及执行框架,这种自进化可提升能力,但成功的变更可能留下无法在不同于变更创建时的状态中安全逆转的持久影响。我们提出EvoUndo,这是一个用于在反事实状态下表示、合成、诊断和独立验证模型生成的自修改可恢复性的框架。在600个未见过的一次性自进化任务中,我们识别出197个可提升能力但未通过可恢复性验证的变更。在原始恢复表示下,传统修复策略在这些自然失败中实现0/197的恢复;确定性预言机分析在原始恢复语言L0下实现48/197的恢复,而扩展恢复演算将经验预言机恢复提升至191/197。随后,一种协议锁定的2×2“按表达力落地”干预分离出两个瓶颈:当原始语言足够时,精确状态地址落地将成功恢复从0/48提升至38/48(79.2%);扩展恢复语言则能在预言机定义的S1层级的143个失败中实现142/143(99.3%)的恢复。在主gpt-oss-120b主干模型上,在更丰富的语言中添加精确地址诊断将恢复降至133/143(93.0%);Qwen3.8-27B复现保留了落地和表达力效应,但未出现这种负向交互,表明后者依赖于模型。这些结果表明,可靠的智能体自进化需要协同设计验证、状态落地、见证语义及恢复语言表达力,而非仅依赖迭代提示。

英文摘要

LLM agents increasingly modify their own prompts, tools, middleware, resources, and execution harnesses at runtime. Such self-evolution can improve capability, but a successful mutation may leave persistent effects that cannot be safely reversed in states different from the one in which it was created. We introduce EvoUndo, a framework for representing, synthesizing, diagnosing, and independently verifying recoverability of model-generated self-modifications across counterfactual states. Across 600 unseen one-shot self-evolution tasks, we identify 197 capability-improving mutations that fail recoverability verification. Under the original recovery representation, conventional repair strategies recover 0/197 of these natural failures. Deterministic oracle analysis recovers 48/197 under the original recovery language L0, while the extended recovery calculus increases empirical oracle recovery to 191/197. A protocol-locked 2x2 grounding-by-expressivity intervention then separates two bottlenecks: exact state-address grounding increases successful recovery from 0/48 to 38/48 (79.2%) when the original language is sufficient, while extending the recovery language enables recovery on 142/143 (99.3%) failures in the oracle-defined S1 stratum. On the primary gpt-oss-120b backbone, adding exact-address diagnostics to the richer language reduces recovery to 133/143 (93.0%); a Qwen3.8-27B replication preserves the grounding and expressivity effects but not this negative interaction, indicating that the latter is model-dependent. These results indicate that reliable agent self-evolution requires co-designing verification, state grounding, witness semantics, and recovery-language expressivity rather than relying on iterative prompting alone.

↑