AI 中文总结
本研究提出动态评估框架ContinualSkillBench,发现大语言模型智能体顺序执行可提升任务性能,上下文学习与显式技能维护效果相当,弱模型易积累零散技能,当前机制仍难整合经验为稳健可迁移技能。
AI 中文摘要
现代智能体框架为大语言模型配备外部技能库以解决复杂任务,但这些系统能否有效发展技能、且所获技能是否能提升任务解决能力仍不明确。为弥合这一差距,我们提出ContinualSkillBench,这是一个用于上下文持续技能学习的动态评估框架,涵盖五个代表性领域,每个领域包含100个相互关联的子任务,按难度递增排序且存在跨任务技能复用机会。我们的实验表明,顺序执行通常会提升性能,但增益因模型和领域差异显著;平均而言,上下文学习与显式技能维护表现相当,说明大部分提升源于对先前上下文和反馈的适应,而非仅可复用技能的抽象;不过显式技能对需要可复用流程或精确输出的任务仍有选择性益处。我们还发现,能力较弱的模型倾向于积累更大、更零散的特定任务技能集合。这些结果表明,当前的上下文技能演化机制可支持持续适应,但仍难以将经验稳定整合为稳健且可迁移的技能。
英文摘要
Modern agent frameworks equip large language models with external skill libraries to solve complex tasks. However, it remains unclear whether these systems can effectively evolve their skills and whether the resulting skills improve task-solving capabilities. To bridge this gap, we introduce ContinualSkillBench, a dynamic evaluation framework for in-context continual skill learning. It covers five representative domains, each containing 100 interconnected subtasks ordered by increasing difficulty and opportunities for cross-task skill reuse. Our experiments show that sequential execution generally improves performance, but the gains vary substantially across models and domains. Moreover, in-context learning performs comparably to explicit skill maintenance on average, suggesting that much of the improvement arises from adaptation to prior context and feedback rather than reusable skill abstraction alone. Explicit skills nevertheless provide selective benefits for tasks requiring reusable procedures or precise outputs. We further find that less capable models tend to accumulate larger, more fragmented collections of task-specific skills. These findings show that current in-context skill evolution mechanisms can support continual adaptation, but still struggle to consistently consolidate experience into robust and transferable skills.