发表机构
Shanghai Jiao Tong University; Zhejiang University; Hithink Research(上海交通大学; 浙江大学; 同花顺研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究智能体能否从自身运行学可复用技能,引入EvoClawBench基准测试,涵盖多任务多子问题并支持多运行时。实验对比不同方式,发现直接基线性能依赖运行时,自我创作技能效果各异,表明学习可复用技能有选择性且成本敏感。
AI 中文摘要
现有智能体基准测试主要关注任务完成、工具使用或技能效用,未考量运行时能否将自身运行证据转化为可复用技能以提升新执行效果。我们引入EvoClawBench,用于重复的、有固定支持任务的闭环技能学习问题基准测试。该基准测试比较无技能直接执行、执行前预技能创作以及首次运行证据的后技能总结及新的第二次执行。套件包含100个任务和502个子问题,支持多种智能体运行时。实验表明直接基线性能强烈依赖运行时,自我创作技能效果不一。这些结果表明从智能体自身运行中学习可复用技能具有选择性且对成本敏感。
英文摘要
Existing agent benchmarks primarily test task completion, tool use, or skill utility, but do not isolate whether a runtime can convert evidence from its own runs into reusable skills that improve fresh executions after authoring overhead. We introduce EvoClawBench, a benchmark for this closed-loop skill-learning question on repeated, fixture-backed tasks. EvoClawBench compares direct execution without skills, PreSkill authoring before execution, and PostSkill summarization from first-run evidence followed by a fresh second execution. The suite contains 100 tasks and 502 sub-problems across coding, data, office, security, operations, and domain-document workflows, with support for multiple agent runtimes. Experiments with OpenClaw and nanobot under local execution show that direct baseline performance is strongly runtime-dependent: OpenClaw remains below 20% across models, while nanobot ranges from 56.45% to 96.13%. Self-authored skills have mixed effects. nanobot GPT-5.4 stays above 96% in all modes and MiniMax-M2.7 improves from 90.97% to 94.50% under PostSkill, but nanobot DeepSeek-V4-Pro drops from 77.77% to 4.80% with PreSkill and 0.99% with PostSkill. OpenClaw shows similarly non-monotonic behavior, with some skill runs near baseline and others collapsing. These results indicate that learning reusable skills from an agent's own runs is selective and cost-sensitive, rather than an automatic benefit of adding skill authoring to an agent loop.