成功留下绕行:为长时程智能体学习可执行的操作指南
Success Leaves Detours: Learning Executable Walkthroughs for Long-Horizon Agents
浏览论文内容
中文总结 AI 辅助
针对稀疏奖励轨迹中失败与绕行问题,提出Trace框架,通过信用引导和依赖切片提取可执行操作指南,在三个基准上显著提升长时程智能体性能。
中文摘要 AI 辅助
测试时自进化智能体通过复用过往经验来提升性能,然而稀疏奖励轨迹中包含失败、循环和绕行,而摘要往往遗漏了执行所需的状态条件和动作依赖。我们研究从稀疏奖励轨迹中归纳可执行的操作指南(Walkthrough):提取紧凑、状态条件化且可验证的程序。我们的关键观察是,延迟信用分配能识别与进展相关的动作,但无法确定这些动作是否产生了后续动作所需的事实。我们提出Trace,一个信用引导、依赖 grounded 的框架,将噪声轨迹编译为可执行的操作指南记忆。它从奖励和持久状态变化中检测进展锚点,传播信用以识别有价值的转换,并从跨回合的成功与失败证据中估计动作前提。随后,反向依赖切片将所需事实追溯到其生产者,提取依赖一致的动作链,同时移除无关循环和绕行。由此产生的操作指南编码了入口条件、有序的状态-动作-效果步骤以及完成和失败谓词,支持复用、中间状态恢复和程序化验证。在J-TTL、WebShop和ScienceWorld上使用三个开源LLM进行的实验表明,Trace持续优于八个测试时学习和记忆基线。与最强基线相比,它分别将平均AUC和Final-$3$提高了$30.0\%$和$40.5\%$,同时使用更少的推理令牌。这些结果表明,长时程交互更多受益于状态条件化的可执行程序,而非完整轨迹或抽象摘要。
英文摘要
Test-time self-evolving agents improve by reusing past experience, yet sparse-reward trajectories contain failures, loops, and detours, while summaries often omit the state conditions and action dependencies needed for execution. We study executable Walkthrough induction from sparse-reward trajectories: extracting compact, state-conditioned, and verifiable procedures. Our key observation is that delayed credit identifies actions associated with progress but cannot determine whether they produce facts required by later actions. We propose Trace, a credit-guided, dependency-grounded framework that compiles noisy trajectories into executable Walkthrough Memory. It detects progress anchors from rewards and persistent state changes, propagates credit to identify valuable transitions, and estimates action prerequisites from cross-episode success and failure evidence. Backward dependency slicing then traces required facts to their producers, extracting dependency-consistent action chains while removing irrelevant loops and detours. The resulting Walkthroughs encode entry conditions, ordered state--action--effect steps, and completion and failure predicates, supporting reuse, intermediate-state resumption, and programmatic verification. Experiments on J-TTL, WebShop, and ScienceWorld with three open-source LLMs show that Trace consistently outperforms eight test-time learning and memory baselines. Compared with the strongest baseline, it improves average AUC and Final-$3$ by $30.0%$ and $40.5%$, respectively, while using fewer inference tokens. These results show that long-horizon interaction benefits more from state-conditioned executable procedures than from complete trajectories or abstract summaries.
发表机构
- Tongji University(同济大学)
- Shanghai AI Laboratory(上海人工智能实验室)
- Fudan University(复旦大学)
机构由 AI 辅助整理,请以论文原文为准。