arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.04008cs.AIcs.LGcs.SE

SkillScriptBench:超越Markdown的可执行智能体技能包自进化基准

SkillScriptBench: Benchmarking Self-Evolution of Executable Agent Skill Packages Beyond Markdown

Yuxuan Liu, Haoran Li, Yuhao Zhang, Jiahe Guo, Hongyu Luo, Wenbin Hu, Huihao Jing, Kawai Chung, Junle Chen, Changxuan Fan, Qing Zong, Lingyun Xie, Yangqiu Song

首次发表
浏览论文内容

中文总结 AI 辅助

提出SkillScriptBench基准,区分文档与脚本修复,并引入AST引导的技能修订方法,显著提升修复成功率与一致性。

中文摘要 AI 辅助

可执行智能体技能将自然语言指令和脚本组合成可复用的包,供LLM智能体使用,修改这些技能需要在不破坏正确行为的前提下修复错误。现有基准在评估技能自进化时,并未系统地区分文档修复、脚本修复和保持原有正确行为。我们引入了SkillScriptBench,一个包含350个任务的基准,旨在分别评估这些能力。通过对超过35,000个托管在GitHub上的技能根目录进行调查,我们选取了100个包并构建了150个修复任务。每个任务将一个包含注入脚本故障的包与一个维护请求及所需行为的可执行检查配对。一个补充的受控轨道包含来自50个包的200个任务,每个任务在四种状态下(干净、文档故障、脚本故障、两者均有故障)以相同的维护请求进行评估。在四个LLM上,同时编辑文档和脚本的方法能够修复脚本故障,但在文档修复或保持正确行为方面并未持续优于仅修改Markdown的方法。因此,我们引入了AST引导的技能修订,该方法利用抽象语法树和调用关系将维护需求关联到相关代码位置。它将脚本编辑限制在这些位置,并更新文档以匹配修订后的脚本。跨模型平均,该修订阶段在故障包上的修复成功率绝对提升为Raw Package的21.9%和CoEvoSkills的27.7%。在所有三次运行中解决任务比例上的绝对提升分别达到20.8%和31.5%,表明在重复运行中修复成功率更加一致。

英文摘要

Executable Agent Skills combine natural-language instructions and scripts into reusable packages for LLM agents, and revising them requires fixing errors without breaking correct behavior. Existing benchmarks do not systematically distinguish documentation repair, script repair, and preservation when evaluating skill self-evolution. We introduce SkillScriptBench, a 350-task benchmark designed to evaluate these capabilities separately. From a survey of over 35,000 GitHub-hosted Skill roots, we select 100 packages and construct 150 repair tasks. Each task pairs a package containing injected script faults with a maintenance request and executable checks of the required behavior. A complementary controlled track contains 200 tasks from 50 packages, each evaluated under the same maintenance request in four states: clean, documentation faults, script faults, and faults in both. Across four LLMs, methods that edit both documentation and scripts can repair script faults but do not consistently outperform Markdown-only revision on documentation repair or preservation. We therefore introduce AST-Guided Skill Revision, which uses abstract syntax trees and calling relationships to link maintenance requirements to relevant code locations. It restricts script edits to these locations and updates the documentation to match the revised scripts. Averaged across models, this revision stage yields absolute gains in repair success of 21.9% for Raw Package and 27.7% for CoEvoSkills on faulty packages. Absolute gains in the proportion of tasks solved in all three runs reach 20.8% and 31.5%, respectively, indicating more consistent repair success across repeated runs.

发表机构

  • The Hong Kong University of Science and Technology(香港科技大学)
  • Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
  • Harbin Institute of Technology, Harbin(哈尔滨工业大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑