arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

回滚世界,保留反思:面向长时程LLM智能体的回滚诱导反思

Rollback the World, Keep the Reflection: Rollback-Induced Reflection for Long-Horizon LLM Agents

Yi Yu, Liuyi Yao, Yaliang Li, Enshu Wang, Libing Wu

arXiv 2609.18304首次发表:更新:

发表机构

Wuhan University; Alibaba Group(武汉大学; 阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对长时程LLM智能体错误累积问题,提出回滚诱导反思(RIR)框架,联合决定干预时机、恢复位置与保留信息,在恢复状态同时保留可复用知识,实验证明其能提升多基准任务性能。

AI 中文摘要

大语言模型(LLM)智能体日益通过多步环境交互来处理长时程任务,然而单个错误动作可能改变后续状态和观察,导致错误随时间累积。现有方法要么在不修复已改变环境状态的情况下纠正上下文,要么在恢复早期状态的同时丢弃有用经验,这使得既难以消除失败条件,又难以避免重复过去的错误。我们认为,可靠的恢复应被视作一个回滚边界控制问题,需要联合决定何时干预、从何处恢复以及哪些信息应在恢复后保留。基于这一观点,我们提出回滚诱导反思(RIR),一个统一的恢复框架,它将执行恢复到选定的先前状态,同时携带从被放弃轨迹中提炼出的可复用知识以指导后续决策。我们进一步通过一个统一算子刻画恢复过程,该算子作用于回滚深度和保留记忆,提供了状态恢复与知识保留的一般视角。在三个长时程基准上的实验表明,RIR在多个LLM骨干网络上持续提升任务性能,其中结构化反思记忆保留了有用经验,选择性回滚实现了高效恢复。

英文摘要

Large language model (LLM) agents increasingly tackle long-horizon tasks through multi-step environment interaction, yet a single erroneous action can alter subsequent states and observations, causing errors to compound over time. Existing methods either correct the context without repairing altered environment states or restore earlier states while discarding useful experience, making it difficult to both eliminate failure conditions and avoid repeating past mistakes. We argue that reliable recovery should instead be treated as a rollback-boundary control problem that jointly determines when to intervene, where to resume, and what information should survive recovery. Based on this view, we propose Rollback-Induced Reflection (RIR), a unified recovery framework that restores execution to a selected prior state while carrying forward reusable knowledge distilled from the abandoned trajectory to guide subsequent decisions. We further characterize recovery through a unified operator over rollback depth and retained memory, providing a general view of state restoration and knowledge retention. Experiments on three long-horizon benchmarks show that RIR consistently improves average task performance across multiple LLM backbones, with structured reflection memory preserving useful experience and selective rollback enabling efficient recovery.

Comments12 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑