arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.11781cs.CV

Skill-V:面向交互智能体的可验证自演进技能库

Skill-V: Verifiable Self-Evolving Skill Library for Interactive Agents

Jie Ma, Zhipeng Qian, Yufei Ma, Zihan Liang, Jiayi Ji, Qingpeng Cai, Ben Chen, Peng Jiang, Xiaoshuai Sun

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出Skill-V可验证自演进技能库,通过带版本可证伪契约、结果驱动演进及适用性过滤,在ALFWorld和WebShop上分别达95.3%、85.9%成功率,实现可靠技能演进。

中文摘要 AI 辅助

交互智能体可将经验转化为可复用技能,但现有自演进技能库主要通过积累新知识实现改进。失败可能催生新技能,而随着新证据的到来,已存储的技能较少被重新审视。然而,仅靠增长无法确保可靠性,因为检索到的技能可能不适用于当前任务条件,且现有技能可能编码了错误指定的操作边界。因此,可靠的技能演进不仅需要添加知识,还需要测试和修订已存储的内容。我们提出Skill-V,一种可验证的自演进技能库。为使存储的知识可测试,我们提出将技能表示为带版本的可证伪契约,该契约将语义意图与可观察的行为标准关联。我们利用环境结果驱动库的演进:具体而言,任务失败促使技能添加,而契约评估与任务结果之间的不一致则指导对现有技能边界的修订。为验证这些修订,我们要求其保留受保护的语义约束,并满足历史重放证据上的评分标准-结果指标的非回归准则。最后,我们采用感知适用性的过滤器,排除被判定为不适用于当前任务的候选技能。在ALFWorld和WebShop上,Skill-V分别实现了95.3%和85.9%的成功率,同时保持了比以增长为导向的基线更紧凑的技能库。感知适用性过滤减少了错误的技能调用,基于结果的修订纠正了错误指定的技能边界,且不会降低对先前观察到的证据的性能。这些结果表明,可靠的技能演进不仅仅需要积累经验:库必须学习保留哪些知识、何时修订知识以及何时应用知识。

英文摘要

Interactive agents can turn experience into reusable skills, yet existing self-evolving skill libraries primarily improve by accumulating new knowledge. Failures may lead to new skills, while previously stored skills are less often revisited as new evidence arrives. However, growth alone does not ensure reliability, as a retrieved skill may be inapplicable under the current task conditions, and an existing skill may encode a mis-specified operational boundary. Reliable skill evolution therefore requires not only adding knowledge, but also testing and revising what is already stored. We introduce Skill-V, a verifiable self-evolving skill library. To make stored knowledge testable, we propose representing skills as versioned, falsifiable contracts that link semantic intent to observable behavioral criteria. We use environment outcomes to drive library evolution. Specifically, task failures motivate skill addition, while disagreements between contract evaluations and task outcomes guide revisions to existing skill boundaries. To validate these revisions, we require them to preserve protected semantic constraints and satisfy non-regression criteria for rubric-outcome metrics on historical replay evidence. Finally, we employ an applicability-aware filter to exclude candidates judged confidently inapplicable to the current task. Across ALFWorld and WebShop, Skill-V achieves success rates of 95.3% and 85.9%, respectively, while maintaining a more compact skill library than growth-oriented baselines. Applicability-aware filtering reduces incorrect skill invocations, and outcome-grounded revisions correct mis-specified skill boundaries without degrading performance on previously observed evidence. These results show that reliable skill evolution requires more than accumulating experience: the library must learn which knowledge to retain, when to revise it, and when it should be applied.

发表机构

  • Xiamen University(厦门大学)
  • Kuaishou Technology(快手科技)
  • Nankai University(南开大学)

机构由 AI 辅助整理,请以论文原文为准。

↑