arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SkillHEX:通过假设驱动的自主探索与利用提升智能体技能

SkillHEX: Improving Agent Skills via Hypothesis-Driven Autonomous Exploration and Exploitation

Yuru Feng, Yaoqi Chen, Beidi Zhao, Qianxi Zhang, Xinjiang Wang, Jianan Lu, Zhirui Wang, Shusen Xu, Zengzhong Li, Qi Chen

arXiv 2608.05628首次发表:更新:

AI 中文总结

SkillHEX是结合假设驱动自验证与证据引导树搜索的闭环框架,在SkillsBench87个任务上,5次迭代预算下用GPT-5.3-Codex和Claude Opus 4.7分别获55.9%、57.9%平均通过率,优于现有自演化方法。

AI 中文摘要

尽管智能体技能为大语言模型(LLMs)提供了可复用的过程性知识,但人工维护存在成本高、可扩展性差以及对齐不当的问题。因此,实际部署需要在测试时进行自主的按需技能演化,且受限于有限的交互预算以及缺少训练或验证集。该场景带来了严重的稀疏奖励挑战,其中结果混杂了多种潜在的失败原因。在这种模糊性下,现有贪婪优化单一现有技能的方法极易陷入利用陷阱,导致早期错误诊断在无成效的轨迹中耗尽有限的试验次数。为解决该问题,我们提出SkillHEX,这是一个结合假设驱动的自验证与证据引导的树搜索的闭环框架。SkillHEX将可证伪的失败假设转化为可执行的测试,生成诊断证据作为密集奖励,无需额外的环境尝试。该证据指导对持久技能修订分支的搜索,动态平衡受支持编辑的利用与合理替代方案的探索。在SkillsBench的87个任务上评估,SkillHEX优于现有自演化方法,在5次迭代预算下,使用GPT-5.3-Codex和Claude Opus 4.7时分别达到55.9%和57.9%的平均通过率。

英文摘要

Although agent skills equip LLMs with reusable procedural knowledge, manual maintenance suffers from high costs, unscalability, and misalignment. Real-world deployments thus require autonomous, on-demand skill evolution at test time, constrained by limited interaction budgets and a lack of training or validation sets. This setting introduces a severe sparse reward challenge, where outcomes conflate multiple latent failure causes. Under such ambiguity, existing methods that greedily refine a single incumbent skill are particularly vulnerable to an exploitation trap, allowing early misdiagnoses to exhaust limited trials along unproductive trajectories. To address this, we introduce SkillHEX, a closed-loop framework coupling hypothesis-driven self-verification with evidence-guided tree search. SkillHEX translates falsifiable failure hypotheses into executable tests, producing diagnostic evidence as dense reward without additional environment attempts. This evidence guides a search over persistent skill-revision branches, dynamically balancing the exploitation of supported edits with the exploration of plausible alternatives. Evaluated on 87 tasks from SkillsBench, SkillHEX outperforms existing self-evolving methods and achieves an average pass rate of 55.9% and 57.9% using GPT-5.3-Codex and Claude Opus 4.7 under a five-iteration budget, respectively.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑