分解与传递:大语言模型智能体的跨任务技能迁移
Break It Down, Pass It On: Cross-Task Skill Transfer in LLM Agents
浏览论文内容
中文总结 AI 辅助
该研究探究LLM智能体跨任务技能迁移的可靠性,对比任务级与子任务级、文本与代码两种技能格式,提出结合特异性与抽象性的技能效用分数,发现子任务级文本技能迁移效果更优,可用于诊断技能记忆。
中文摘要 AI 辅助
大语言模型(LLM)智能体可从已完成的任务中归纳技能,并在后续复用这些技能,从而随经验增长提升能力。但在实际应用中,归纳出的技能迁移可能不可靠,甚至会对检索该技能的智能体造成损害。智能体归纳的技能何时能跨任务可靠迁移,仍是一个未解决的问题。我们开展了一项全面且受控的研究,探究技能归纳方式如何影响其跨任务迁移效果,具体对比了现有方法存在差异的两个维度:任务级与子任务级技能归纳,以及文本与代码两种技能格式。结果显示,任务级技能大多会使智能体的性能降至无记忆基准以下,而子任务级技能平均而言会使性能高于基准;此外,文本格式技能的迁移效果优于代码格式。为进一步解释这些发现,我们考察了归纳技能的两个互补属性:特异性(衡量技能与实际任务的匹配紧密程度)和抽象性(衡量技能的相关性在不同任务间的分布均匀性)。单一属性无法预测任务成功,但二者的组合效应可以,我们将其命名为技能效用分数。该分数在技能迁移时与任务成功度始终相关,且子任务级和文本格式技能的分数更高。计算技能效用仅需技能和任务描述,无需任何任务执行过程,因此我们的分数可在新任务运行前作为技能记忆的实用诊断工具。
英文摘要
Large language model (LLM) agents can induce skills from completed tasks and reuse them later to grow more capable with experience. In practice, induced skills may transfer unreliably and can even harm the agent that retrieves them. When agent-induced skills transfer reliably across tasks remains an open question. We conduct a comprehensive and controlled study of how the way skills are induced shapes their transfer across tasks. Specifically, we compare task-level with subtask-level skill induction and text with code skill formats, the two axes along which existing methods differ. Task-level skills mostly reduce the agent's performance below its no-memory baseline while subtask-level skills raise it above on average, and text skills transfer better than code skills. To further understand our findings, we examine two complementary properties of the induced skills: specificity, which measures how closely a skill matches real tasks, and abstractness, which measures how evenly its relevance spreads across tasks. Neither property alone predicts task success, but their combined effect does, which we propose as a skill utility score. The score correlates consistently with task success when skills are transferred, and subtask-level and text skills score higher. Computing skill utility only needs the skills and task descriptions but not any task execution, so our score serves as a practical diagnostic of a skill memory before any new task runs.
发表机构
- Stony Brook University(石溪大学)
机构由 AI 辅助整理,请以论文原文为准。