arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37128cs.AI

SkillCome:基于双记忆的组对比技能优化

SkillCome: Group Contrast Skill Optimization with Dual Memory

发表机构复旦大学 · 蚂蚁集团 · 上海交通大学
查看机构详情
  • Fudan University(复旦大学)
  • Ant Group(蚂蚁集团)
  • Shanghai Jiao Tong University(上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

Haolin Li, Feng Hong, Ang Li, Chilin Fu, Weichang Wu, Ya Zhang, Yanfeng Wang, Xiaolu Zhang, Jiangchao Yao

首次发表
浏览论文内容

中文总结 AI 辅助

SkillCome提出基于组对比优化与双记忆的技能进化方法,通过多轨迹对比和双记忆积累,提升大型语言模型在问答、推理及智能体任务中的性能,在六个基准上最高提升+5.69个百分点。

中文摘要 AI 辅助

技能进化通过分析在给定技能下生成的轨迹并相应修改技能来提升大型语言模型的能力。现有方法通常每个问题仅生成一条轨迹。然而,这提供的优化信号不足,因为需要从单一轨迹中推断出有效的技能编辑。在失败的轨迹中难以确定是哪些动作导致了失败,或者在成功的轨迹中难以确定哪些动作应被纳入技能。此外,它们依赖局部轨迹批次进行分析,使得优化方向容易受到噪声证据的影响。为解决这些问题,我们提出了SkillCome,一种基于组对比优化与双记忆的技能进化方法。对于每个问题,SkillCome生成多条轨迹并进行组对比分析,以精确识别成功与失败轨迹之间的关键行为差异,从而提供可靠的优化信号。双记忆系统进一步从历史步骤中积累证据,以追踪不同组间共享的模式,从而产生更泛化的优化方向。综合来看,SkillCome构建了一个系统化的优化过程,将观察到的成功轨迹中的经验转化为可复用的技能。在涵盖问答、推理和智能体任务的六个基准上的大量实验证明了我们方法的有效性。SkillCome在五个不同家族和规模的模型上 consistently 优于基线,最高提升达+5.69个百分点。

英文摘要

Skill evolution improves the capabilities of large language models by analyzing trajectories generated under a given skill and modifying the skill accordingly. Existing approaches typically generate a single trajectory per question. However, this provides insufficient optimization signals since it requires inferring effective skill edits from a solitary path. It is difficult to pinpoint which actions caused the failure in a failed trajectory, or to determine which actions in a successful one should be incorporated into the skill. Furthermore, they rely on a local batch of trajectories for analysis, making the optimization direction susceptible to noisy evidence. To address these, we propose SkillCome, a Skill-evolution method based on group Contrast optimization with dual memory. For each question, SkillCome generates trajectories and performs group contrast analysis to precisely identify key behavioral divergences between successful and failed trajectories, offering reliable optimization signals. The dual memory system further accumulates evidence from historical steps to track patterns shared across different groups, leading to more generalized optimization directions. Together, SkillCome builds a systematic optimization process that transforms experience from observed successful trajectories into reusable skills. Extensive experiments on six benchmarks spanning question answering, reasoning, and agentic tasks demonstrate the effectiveness of our method. SkillCome consistently outperforms baselines across five models of varying families and scales, with gains up to +5.69 points.

补充信息

↑