arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SkillFocus:通过能力分解演化智能体技能

SkillFocus: Evolving Agent Skills via Capability Decomposition

Ning Wang, Zhiren Gong, Bingdong Li, Peng Yang, Aimin Zhou

arXiv 2609.34397首次发表:更新:

发表机构

East China Normal University; Nanyang Technological University; Southern University of Science and Technology; Shanghai Innovation Institute(华东师范大学; 南洋理工大学; 南方科技大学; 上海创新研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SkillFocus通过将任务需求分解为固定能力空间,识别未解决任务对应的能力并指导修订,在四个基准上平均超越最强基线5.7个百分点,同时减少24%的演化令牌。

AI 中文摘要

智能体技能演化旨在通过迭代修订来改进大型语言模型(LLM)智能体的可重用程序性指导。现有方法主要基于执行轨迹或反馈进行每次修订,使得跨任务反复出现的行为需求隐含化,并将修订与当前技能的行为绑定。我们提出SkillFocus,它将反复出现的任务需求分解到一个能力空间,该空间在技能演化过程中保持固定,从而将任务需求与当前技能的行为方式分离开来。SkillFocus将当前任务结果映射到该空间,以识别导致最多任务未解决的能力,然后利用该能力来确定修订内容以及使用哪些证据。在涵盖异构任务的四个基准上,SkillFocus在所有四个基准上均取得了最佳的保留准确率,平均比最强竞争结果高出5.7个百分点,同时平均比最接近的迭代基线少使用24%的演化令牌。受控研究进一步表明,源自反复出现的任务需求的能力优于基于任务语义和基于执行轨迹的替代方案,而随机化任务-能力分配会使最终准确率降低多达20.2个百分点。在优先修订下,将证据与所选能力匹配可使候选增益提高4.4个百分点。

英文摘要

Agent skill evolution seeks to improve reusable procedural guidance for large language model (LLM) agents through iterative revision. Existing methods base each revision mainly on execution trajectories or feedback, leaving recurring behavioral requirements across tasks implicit and tying revision to the behavior of the current skill. We introduce SkillFocus, which decomposes recurring task requirements into a capability space that remains fixed as the skill evolves, separating what tasks require from how the current skill behaves. SkillFocus maps current task outcomes to this space to identify the capability that leaves the most tasks unresolved, then uses that capability to determine what to revise and which evidence to use. Across four benchmarks spanning heterogeneous tasks, SkillFocus achieves the best held-out accuracy on all four, outperforming the strongest competing result by 5.7 points on average while using 24\% fewer evolution tokens on average than the closest iterative baseline. Controlled studies further show that capabilities derived from recurring task requirements outperform task-semantic and execution-derived alternatives, while randomizing task--capability assignments reduces final accuracy by up to 20.2 points. Matching evidence to the selected capability increases candidate gain by 4.4 points under prioritized revision.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑