当合成数据有害:LLM智能体技能检索中的灾难性遗忘
When Synthetic Data Hurts: On Catastrophic Forgetting in Skill Retrieval for LLM Agents
浏览论文内容
中文总结 AI 辅助
本研究针对LLM智能体技能检索中合成数据微调导致的灾难性遗忘问题,提出并评估了多种持续学习方法,在保持OOD性能的同时提升分布内检索13.98%。
中文摘要 AI 辅助
LLM智能体日益依赖运行时检索到的外部技能,这使得从大型技能库中进行技能选择成为一个关键挑战。我们提出了一个覆盖34,396个技能的生产级技能路由器,并开展了一项使用有限真实监督和合成数据的大规模技能检索研究。我们发现,合成数据微调虽然改善了分布内检索,但会导致在真实数据和分布外(OOD)数据上的灾难性遗忘。我们评估了多种受持续学习启发的缓解遗忘的微调方法,包括嵌入锚定正则化、无遗忘学习(LwF)、弹性权重巩固(EWC)和L2初始化。结果表明,这些方法不仅保持了在OOD技能检索上的性能,还使0.6B Qwen检索器和重排序器在合成分布内技能上的检索性能提升了13.98%。我们的结果为稀缺、多正样本监督提供了一个实用的基准和稳健的微调方案。
英文摘要
LLM agents increasingly rely on external skills retrieved at runtime, making skill selection from large repositories a critical challenge. We present a production skill router over 34,396 skills and a large-scale study of skill retrieval using limited real supervision and synthetic data. We found that the synthetic-data fine-tuning improves in-distribution retrieval but it causes catastrophic forgetting on real and out-of-distribution (OOD) data. We evaluate several forgetting mitigation fine-tuning approaches inspired by continual learning, including embedding-anchor regularization, Learning without Forgetting (LwF), Elastic Weight Consolidation (EWC), and L2-initialization. The results show that these approaches not only retain the performance on OOD skills retrieval but also improve the retrieval on synthetic in-distribution skills by 13.98\% for 0.6B Qwen retriever and reranker. Our results provide a practical benchmark and a robust fine-tuning recipe for scarce, multi-positive supervision.
发表机构
- Manulife(宏利金融)
机构由 AI 辅助整理,请以论文原文为准。