arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15165cs.AI

SkillCommit:通过行为验证的范围扩展实现智能体技能演化

SkillCommit: Evolving Agent Skills through Behaviorally Validated Scope Expansion

Yu He, Weikai Yang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出SkillCommit框架,通过保留经验的验证行为、抽象相关技能,在RuleArena等数据集上提升智能体性能,且所学技能可跨模型迁移。

中文摘要 AI 辅助

大型语言模型(LLM)智能体无需参数更新即可通过将历史经验转化为可复用的过程知识实现持续改进。然而,现有方法通常基于语义相似性或LLM判断来整合经验,这可能会合并表面相关但行为不兼容的策略,进而降低性能。为解决该问题,我们提出SkillCommit,这是一种在线技能演化框架,可不断将经验转化为分层的可复用技能库。每个新经验最初会被保存为特定实例的补丁,保留其局部上下文中验证过的行为。随着相关技能的积累,SkillCommit会将那些共享共同行为机制的技能抽象为更高层级的技能。具体而言,对于每个传入的技能,基于嵌入的检索首先识别候选相关技能;跨实例重放和基于LLM的机制检查确定这些技能是否可跨案例迁移并共享共同的底层机制;通过两项检查的候选会被抽象为更高层级的技能,且仅当该技能保留所有构成技能的验证行为时才会被提交。在RuleArena、OpenExempt和KOR-Bench上的实验表明,SkillCommit在不同领域均能持续提升智能体性能,此外,所学技能可跨模型规模和模型族迁移,实现跨模型经验迁移。

英文摘要

Large language model (LLM) agents can continually improve without parameter updates by converting historical experience into reusable procedural knowledge. However, existing methods often consolidate experience based on semantic similarity or LLM judgments, which may merge superficially related but behaviorally incompatible strategies and thereby degrade performance. To address the issue, we propose SkillCommit, an online skill evolution framework that continuously transforms experience into a hierarchical library of reusable skills. Each new experience is initially preserved as an instance-specific patch, retaining the behavior validated in its local context. As related skills accumulate, SkillCommit abstracts those sharing a common behavioral mechanism into higher-level skills. Specifically, for each incoming skill, embedding-based retrieval first identifies candidate related skills. Cross-instance replay and an LLM-based mechanism check determine whether these skills transfer across cases and share a common underlying mechanism. Candidates that pass both checks are abstracted into a higher-level skill and committed only if it preserves the validated behavior of all constituent skills. Experiments on RuleArena, OpenExempt and KOR-Bench demonstrate that SkillCommit consistently improves agent performance across diverse domains. Moreover, the learned skills transfer across model scales and families, enabling cross-model experience transfer.

↑