arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10750cs.IRcs.AIcs.LG

当合成数据有害:LLM智能体技能检索中的灾难性遗忘

When Synthetic Data Hurts: On Catastrophic Forgetting in Skill Retrieval for LLM Agents

Syed Shariyar Murtaza, Yifan Nie, Utkarsh Soni, Eugene Wen, Arvid Frydenlund

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对LLM智能体技能检索中合成数据微调导致的灾难性遗忘问题,提出并评估了多种持续学习方法,在保持OOD性能的同时提升分布内检索13.98%。

中文摘要 AI 辅助

LLM智能体日益依赖运行时检索到的外部技能,这使得从大型技能库中进行技能选择成为一个关键挑战。我们提出了一个覆盖34,396个技能的生产级技能路由器,并开展了一项使用有限真实监督和合成数据的大规模技能检索研究。我们发现,合成数据微调虽然改善了分布内检索,但会导致在真实数据和分布外(OOD)数据上的灾难性遗忘。我们评估了多种受持续学习启发的缓解遗忘的微调方法,包括嵌入锚定正则化、无遗忘学习(LwF)、弹性权重巩固(EWC)和L2初始化。结果表明,这些方法不仅保持了在OOD技能检索上的性能,还使0.6B Qwen检索器和重排序器在合成分布内技能上的检索性能提升了13.98%。我们的结果为稀缺、多正样本监督提供了一个实用的基准和稳健的微调方案。

英文摘要

LLM agents increasingly rely on external skills retrieved at runtime, making skill selection from large repositories a critical challenge. We present a production skill router over 34,396 skills and a large-scale study of skill retrieval using limited real supervision and synthetic data. We found that the synthetic-data fine-tuning improves in-distribution retrieval but it causes catastrophic forgetting on real and out-of-distribution (OOD) data. We evaluate several forgetting mitigation fine-tuning approaches inspired by continual learning, including embedding-anchor regularization, Learning without Forgetting (LwF), Elastic Weight Consolidation (EWC), and L2-initialization. The results show that these approaches not only retain the performance on OOD skills retrieval but also improve the retrieval on synthetic in-distribution skills by 13.98\% for 0.6B Qwen retriever and reranker. Our results provide a practical benchmark and a robust fine-tuning recipe for scarce, multi-positive supervision.

发表机构

  • Manulife(宏利金融)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑