发表机构
SANKEN, Osaka University(大阪大学产业科学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究探讨自我进化技能能否泛化到新任务,发现多数技能仅部分迁移,并提出GSO方法,通过为每任务重写技能,在六个基准上取得最优表现。
AI 中文摘要
AI智能体可以将从过去任务中学到的内容外部化为可复用的技能,例如程序、检查清单、代码或其他可执行的工件,这些技能在解决新任务时可以被检索和复用。自我进化技能方法在训练任务上每轮实践后不断重写这些技能,然后该技能被用于同类的新任务。我们提出一个问题:一项技能在其训练任务上表现出的改进能否迁移到新的测试任务上?我们在六个基准上测试了五种自我进化方法和一种一次性技能,所有方法使用相同的模型、相同的智能体和相同的训练/测试划分。在21项在训练任务上有所改进的技能中,有5项在测试任务上保留了全部改进,13项保留了部分改进,3项没有保留任何改进。没有一种现有方法在所有场景中都是最优的。当我们阅读这些技能时,那些迁移效果差的技能往往修复了应该依赖于任务的细节,例如列名和输出文件,或者将对一次失败的修复变成了适用于所有任务的规则。一个阅读技能内容的LLM评判者通常能看到这一点:它对完成技能的排序与测试结果在86%的配对中一致。但它对单次编辑效果的预测很差,因此编辑仍然需要通过运行来测试。基于这些发现,我们描述了通用技能优化(GSO),它只保留编写技能的指南,并为每个任务编写新技能;它在所有六个基准上得分最高。
英文摘要
AI agents can externalize what they learn from past tasks into reusable \emph{skills}, such as procedures, checklists, code, or other executable artifacts, that can be retrieved and reused when solving new tasks. Self-evolving skill methods keep rewriting these skills after each round of practice on training tasks, and the skill is then used on new tasks of the same kind. We ask a question: does the improvement a skill shows on its training tasks carry over to new test tasks? We test five self-evolving methods and a one-shot skill on six benchmarks, with the same model, the same agent, and the same train/test split for every method. Of the 21 skills that improve on their training tasks, 5 keep all of that improvement on the test tasks, 13 keep part of it, and 3 keep none of it. No existing method is best everywhere. When we read the skills, the ones that carry over badly often fix details that should depend on the task, such as column names and output files, or turn a fix for one failure into a rule for every task. An LLM judge that reads the skill content can often see this: it ranks finished skills the same way the test results do in 86\% of pairs. But it predicts the effect of a single edit poorly, so edits still have to be tested by running them. Based on these findings, we describe Generalizable Skill Optimization (GSO), which keeps only a guide for writing skills and writes a new skill for each task; it scores highest on all six benchmarks.