AI 中文总结
SkillHEX是结合假设驱动自验证与证据引导树搜索的闭环框架,在SkillsBench87个任务上,5次迭代预算下用GPT-5.3-Codex和Claude Opus 4.7分别获55.9%、57.9%平均通过率,优于现有自演化方法。
AI 中文摘要
尽管智能体技能为大语言模型(LLMs)提供了可复用的过程性知识,但人工维护存在成本高、可扩展性差以及对齐不当的问题。因此,实际部署需要在测试时进行自主的按需技能演化,且受限于有限的交互预算以及缺少训练或验证集。该场景带来了严重的稀疏奖励挑战,其中结果混杂了多种潜在的失败原因。在这种模糊性下,现有贪婪优化单一现有技能的方法极易陷入利用陷阱,导致早期错误诊断在无成效的轨迹中耗尽有限的试验次数。为解决该问题,我们提出SkillHEX,这是一个结合假设驱动的自验证与证据引导的树搜索的闭环框架。SkillHEX将可证伪的失败假设转化为可执行的测试,生成诊断证据作为密集奖励,无需额外的环境尝试。该证据指导对持久技能修订分支的搜索,动态平衡受支持编辑的利用与合理替代方案的探索。在SkillsBench的87个任务上评估,SkillHEX优于现有自演化方法,在5次迭代预算下,使用GPT-5.3-Codex和Claude Opus 4.7时分别达到55.9%和57.9%的平均通过率。
英文摘要
Although agent skills equip LLMs with reusable procedural knowledge, manual maintenance suffers from high costs, unscalability, and misalignment. Real-world deployments thus require autonomous, on-demand skill evolution at test time, constrained by limited interaction budgets and a lack of training or validation sets. This setting introduces a severe sparse reward challenge, where outcomes conflate multiple latent failure causes. Under such ambiguity, existing methods that greedily refine a single incumbent skill are particularly vulnerable to an exploitation trap, allowing early misdiagnoses to exhaust limited trials along unproductive trajectories. To address this, we introduce SkillHEX, a closed-loop framework coupling hypothesis-driven self-verification with evidence-guided tree search. SkillHEX translates falsifiable failure hypotheses into executable tests, producing diagnostic evidence as dense reward without additional environment attempts. This evidence guides a search over persistent skill-revision branches, dynamically balancing the exploitation of supported edits with the exploration of plausible alternatives. Evaluated on 87 tasks from SkillsBench, SkillHEX outperforms existing self-evolving methods and achieves an average pass rate of 55.9% and 57.9% using GPT-5.3-Codex and Claude Opus 4.7 under a five-iteration budget, respectively.