AI 中文总结
研究现实世界中智能体学习受环境交互限制的问题,提出“经验蒸馏”,开发无需额外环境交互的实现方法,实验表明该方法能高效保留上下文学习收益,相比经典方法减少环境样本使用。
AI 中文摘要
现实世界中的智能体学习常常受到昂贵环境交互的限制,如运行耗时实验或获取人类反馈。上下文学习为智能体从自身交互历史中学习提供了高效样本方式,但经验移除后收益消失。上下文蒸馏能将上下文信息内化到模型权重中,然而在不牺牲环境样本效率的情况下应用于智能体交互历史的研究尚少。我们提出“经验蒸馏”问题并开发了一种实现方法,实验表明该方法在两个领域至少保留了64.8%的上下文学习收益,远超直接监督微调的3.8%,且与经典强化学习基线相比,所需环境样本至少少9.6倍。
英文摘要
Real-world agent learning is often constrained by costly environment interactions, such as running time-consuming experiments or obtaining human feedback. In-context learning offers a highly sample-efficient way for agents to learn from their own interaction histories, but its gains disappear once that experience is removed from the context. Separately, context distillation provides a mechanism for internalizing contextual information into model weights. However, applying it to agents' interaction histories without sacrificing environment sample efficiency remains underexplored. We term this problem Experience Distillation and develop an implementation that requires no further environment interaction beyond the collected experience. Experiments on 749 curated software-engineering tasks and six text-adventure games show that it retains at least 64.8\% of the gains from in-context learning across both domains, whereas direct supervised fine-tuning on the collected experience recovers only 3.8\%. Compared with classical reinforcement-learning baselines, in-context learning from trial-and-error experience followed by Experience Distillation matches their performance with at least \(9.6\times\) fewer environment samples.