发表机构
School of Artificial Intelligence, University of Chinese Academy of Sciences; Institute of Automation, Chinese Academy of Sciences; School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences; Wuhan AI Research(中国科学院大学人工智能学院; 中国科学院自动化研究所; 中国科学院大学先进交叉科学学院; 武汉人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出EDGE框架,将检索的经验内化到策略中,在ALFWorld和WebShop数据集7B规模下较GRPO提升8.3、12.5个成功率百分点,移除外部经验后仍保留96.0%支架性能。
AI 中文摘要
基于结果目标的强化学习(如GRPO)可使基于大语言模型(LLM)的智能体解决复杂、长程任务,但交互轨迹中嵌入的可复用探索模式在单次策略更新后大多被丢弃。现有经验增强方法在推理时检索历史指导,但应用经验时未考虑策略的 evolving 能力,且会产生对外部检索的持续依赖。本文提出EDGE(Experience-Distillation for Guided Exploration)框架,将检索到的经验视为临时训练阶段的支架,并逐步将其益处内化到参数化策略中。具体而言,EDGE将每个回合组划分为经验条件轨迹和无经验轨迹,以估计并仅接纳正边际增益,无需额外采样;随后通过自身经验支持上的反向KL目标,将诱导行为蒸馏到基础策略中。协同进化经验库还会在策略演化时合成来自新出现失败模式的指导,并修剪过时条目。在ALFWorld和WebShop数据集上,EDGE在7B规模下较GRPO分别提升8.3和12.5个成功率百分点,且在推理时移除外部经验后仍保留其支架性能的96.0%。代码可在该https URL获取。
英文摘要
Reinforcement learning with outcome-based objectives such as GRPO enables LLM-based agents to solve complex, long-horizon tasks, yet the reusable exploration patterns embedded in interaction trajectories are largely discarded after a single policy update. Existing experience-augmented approaches retrieve historical guidance at inference time, but they apply experiences without accounting for the policy's evolving capability and create persistent dependencies on external retrieval. We propose EDGE (Experience-Distillation for Guided Exploration), a framework that treats retrieved experiences as temporary training-time scaffolds and progressively internalizes their benefits into the parametric policy. Concretely, EDGE partitions each rollout group into experience-conditioned and experience-free trajectories to estimate and admit only positive marginal gains without extra sampling, then distills the induced behavior into the base policy via a reverse-KL objective on its own empirical support. A co-evolutionary experience bank further synthesizes guidance from emerging failure modes and prunes obsolete entries as the policy evolves. Across embodied, web, and search-based QA tasks, EDGE improves over strong RL baselines by up to 12.5 points and remains effective without inference-time scaffolds or a proprietary reflector. The code is available at https://github.com/xvolcano02/EDGE.
CommentsAccepted to EMNLP 2026 (Main Conference)