arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14380cs.AI

AgentRewind:面向长 horizon LLM 智能体的可恢复执行框架

AgentRewind: Recoverable Execution for Long-Horizon LLM Agents

Yu Zhuang, Kefei Chen, Yitong Duan, Shuxin Zheng, Jian Li, Xu-Yao Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

AgentRewind 是面向长周期 LLM 智能体的运行时恢复框架,通过记录检查点支持状态回退,结合自建基准 MettleBench 实验,可提升智能体任务成功率与清单进度。

中文摘要 AI 辅助

许多现实任务要求 LLM 智能体在长执行周期中与环境交互,执行早期出现的错误可能会通过智能体上下文和环境状态传播,其影响难以通过后续操作逆转。现有方法主要通过计划优化和安全检查减少此类错误,但在错误发生后几乎无法提供支持。为实现长周期执行中的恢复,我们提出 AgentRewind,这是一种运行时恢复框架,可记录智能体上下文和受控环境的对齐检查点,允许智能体返回更早的状态并利用之前尝试的信息恢复执行。我们还构建了 MettleBench,这是一个用于评估长周期工程任务的任务完成度和部分进展的基准,这类任务包含一系列相关需求。在跨任务、多模型、执行策略和智能体框架的实验中,结果显示 AgentRewind 相比基线方法提升了任务成功率和平均清单进度。

英文摘要

Many real-world tasks require LLM agents to interact with their environments over long execution horizons. Errors that occur early in execution may propagate through both the agent context and environment state, and their effects may be difficult to reverse through subsequent actions. Existing methods mainly seek to reduce such errors through plan refinement and safety checks but provide little support after errors occur. To enable recovery during long-horizon execution, we present AgentRewind, a runtime recovery framework that records aligned checkpoints of the agent context and controlled environment, allowing agents to return to an earlier state and resume execution with information from previous attempts. We also construct MettleBench, a benchmark for evaluating task completion and partial progress on long-horizon engineering assignments containing a series of related requirements. Experiments across tasks, multiple models, execution strategies, and agent harnesses show that AgentRewind improves task success rate and average checklist progress over the compared baselines.

补充信息

↑