发表机构
Zhejiang University; National University of Singapore; Shanghai Jiao Tong University; Meituan(浙江大学; 新加坡国立大学; 上海交通大学; 美团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SkillRise是跨任务学习技能的统一强化学习框架,解耦信用分配机制实现任务求解与技能整理,在多基准上Pass@1性能优于基线,还降低了运行开销。
AI 中文摘要
大型语言模型智能体常遇到相关但不同、共享可复用解决方案模式的任务,但标准智能体强化学习将任务视为独立回合,现有技能学习方法要么聚焦单一任务的重复尝试,要么使用包含提取、检索、执行多阶段的流水线,使各环节相互纠缠。我们提出SkillRise,一种用于跨任务学习技能的统一强化学习框架。SkillRise将相关实例组织成逐步提升难度的序列,使用单一策略在任务求解与整理演化技能文档间交替,该文档会直接传递给下一个任务。跨任务的解耦信用分配机制,既用当前任务结果监督求解过程,又用折扣后的下游结果监督整理过程。在ALFWorld、WebShop和ScienceWorld上的实验显示,SkillRise在对比方法中实现了最强的Pass@1性能,相较于最强基线的提升幅度为2.3至8.5个百分点。尽管在不同任务上训练,其学习到的整理策略对同一任务的重复尝试仍有效。进一步分析揭示了跨任务测试时的缩放效应:即使每个任务仅尝试一次,性能也会随相关任务序列变长而提升,该趋势表明SkillRise在跨任务复用可迁移技能,而非得益于同一任务的重复采样。SkillRise还在大幅降低多阶段技能学习流水线运行开销的同时,保持了强性能。综上,这些结果为大型语言模型智能体提供了一种简单高效的训练范式,使其能在跨任务中提取、优化并复用可迁移技能。
英文摘要
Large language model agents often encounter related yet distinct tasks that share reusable solution patterns. Yet standard agentic reinforcement learning treats tasks as independent episodes, while existing approaches to skill learning either focus on repeated attempts of one task or use pipelines with multiple stages that entangle extraction, retrieval, and execution. We introduce SkillRise, a unified reinforcement learning framework for learning skills across tasks. SkillRise organizes related instances into progressively challenging sequences and uses a single policy to alternate between task solving and curating an evolving skill document passed directly to the next task. Decoupled credit assignment across tasks supervises solving with the current task outcome and curation with discounted downstream outcomes. Experiments on ALFWorld, WebShop, and ScienceWorld show that SkillRise achieves the strongest Pass@1 performance among the compared methods, with gains over the strongest baseline ranging from 2.3 to 8.5 percentage points. Although trained across distinct tasks, its learned curation policy remains effective for repeated attempts on the same task. Further analysis reveals scaling at test time across tasks: performance improves with longer sequences of related tasks even when each task is attempted only once. This trend suggests that SkillRise reuses transferable skills across tasks rather than benefiting from repeated sampling of the same task. SkillRise further retains strong performance while substantially reducing the runtime overhead of skill learning pipelines with multiple stages. Together, these results provide a simple and efficient training paradigm for LLM agents to extract, refine, and reuse transferable skills across tasks.