arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02356cs.AI

SkillTrace:遍历查询-技能图以构建可组合的大语言模型智能体

SkillTrace: Traversing a Query-Skill Graph for Composable LLM Agents

Yue Yao, Shengyuan Wang, Xin Chen, Minke Zhang, Jia He, Bingjun Luo, Tom Gedeon

AI总结:

SkillTrace通过查询-技能图的三个层级关系实现技能组合,在SkillsBench和ALFWorld上取得SOTA性能,且对不同骨干语言模型具备通用性。

AI中文摘要:

大型语言模型智能体越来越多地通过从库中组合可复用技能来解决复杂任务。要实现这一点,关键挑战不仅在于检索单独相关的技能,还在于识别完整且可执行的技能组合。本文认为该问题可在包含三个层级的图中解决:技能查询间的组合关系、查询与技能库中候选技能的相似性,以及所选候选技能间的依赖关系。我们提出SkillTrace,它将用户查询组织成语义层次结构,匹配技能查询与候选技能,并沿技能依赖关系传播。在SkillsBench和ALFWorld上的实验表明,SkillTrace实现了SOTA性能,在SkillsBench上的成功率达53.17%,在ALFWorld上达91.43%;其在不同骨干语言模型上均实现稳定提升,证明了基于图的技能检索的通用性与鲁棒性。

英文摘要:

Large language model agents increasingly solve complex tasks by composing reusable skills from a library. To address this, the key challenge is not merely to retrieve individually relevant skills, but to identify a complete and executable skill composition. In this paper, we argue that this problem can be solved in a graph with three levels: compositional relations among skill queries, similarity between queries and candidates in the skill library, and the dependencies among the selected candidates. We introduce SkillTrace, which organizes the user query into a semantic hierarchy, matches skill queries and candidates, and propagates over the skill dependencies. Experiments on SkillsBench and ALFWorld demonstrate that SkillTrace achieves state-of-the-art performance, reaching a success rate of 53.17% on SkillsBench and 91.43% on ALFWorld. SkillTrace also delivers consistent improvements across different backbone language models, demonstrating the generality and robustness of graph-based skill retrieval.

↑