发表机构
Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学高瓴人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对迭代式智能体自蒸馏中的性能崩溃,提出ReSAIL方法,通过选择性蒸馏和保留PI条件行为,在ALFWorld和TextCraft上实现最终轮成功率平均绝对提升22.5%。
AI 中文摘要
迭代式自蒸馏使大语言模型(LLM)智能体能够从连续部署中学习,为递归式自我改进(RSI)提供了一条路径。然而,我们使用现有方法进行的实验揭示了各轮次间部署性能的崩溃,同时,具有特权信息(PI)的任务性能也在下降。我们通过优先蒸馏信息丰富的交互步骤,并在学生成为下一轮教师时保留PI条件行为来解决这一崩溃问题。我们提出了用于迭代式自蒸馏的保留性与选择性增强(ReSAIL),这是一种用于基于PI的迭代式自蒸馏的即插即用增强方法。ReSAIL选择PI对教师预测改变最大的交互步骤,并在各轨迹间平衡由此产生的蒸馏损失。它还将学生的PI条件输出分布向冻结教师在选定和未选定步骤上的分布进行正则化,以保留PI条件行为,用于下一轮的监督。在ALFWorld和TextCraft上,ReSAIL在三轮中跨模型规模保持了显著收益,在添加到自蒸馏基线时,最终轮成功率平均绝对提升了22.5%。敏感性引导的离线数据选择还提高了AITZ上多模态GUI智能体的动作预测准确性。这些发现首次证明,一种更稳健的学习机制可以有效缓解迭代式智能体自蒸馏在部署轨迹上的性能崩溃。
英文摘要
Iterative self-distillation enables LLM agents to learn from successive deployments, offering a path toward recursive self-improvement (RSI). Yet our experiments with existing methods reveal a collapse in deployment performance across cycles, while task performance with privileged information (PI) also declines. We address this collapse by prioritizing informative interaction steps for distillation and preserving PI-conditioned behavior as the student becomes the next teacher. We introduce Retentive and Selective Augmentation for Iterative Self-Distillation (ReSAIL), a plug-in augmentation for iterative PI-based self-distillation. ReSAIL selects interaction steps where PI most strongly changes the teacher's predictions and balances the resulting distillation losses across trajectories. It also regularizes the student's PI-conditioned output distributions toward those of the frozen teacher at selected and unselected steps to preserve PI-conditioned behavior for supervision in the next cycle. On ALFWorld and TextCraft, ReSAIL sustains substantial gains across model scales over three cycles, with an average absolute gain of 22.5% in final-cycle success rates when added to self-distillation baselines. Sensitivity-guided selection of offline data also improves action prediction accuracy for multimodal GUI agents on AITZ. These findings provide the first evidence that a more robust learning mechanism can effectively mitigate performance collapse in iterative agent self-distillation over deployment trajectories.