CaSKG:用于可扩展智能体技能检索的反事实因果技能图
CaSKG: Counterfactual-Causal Skill Graphs for Scalable Agent Skill Retrieval
- School of Artificial Intelligence, Jilin University(吉林大学人工智能学院)
- Ant Group(蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本研究提出CaSKG框架,通过构建并校准反事实因果技能图实现可扩展智能体技能检索,在6种LLM骨干模型的两个基准测试中均取得最高任务得分,显著优于Graph-of-Skills。
中文摘要 AI 辅助
可复用技能库使大语言模型(LLM)智能体能够在不同任务间复用过程性知识,但也将内存访问转化为具有挑战性的检索问题。全库提示法虽能保持覆盖范围,但上下文成本高;向量检索返回紧凑邻域,却将技能视为独立文本;基于图的检索仅在承载相关性的边可靠时才能恢复工作流上下文。我们提出CaSKG,一种在检索前校准过程关系的反事实因果技能图框架。CaSKG首先从语义、词汇、输入/输出及结构证据构建高召回率有向候选图,利用修复证据和可选LLM判断进一步优化候选得分;随后应用方向条件文本反事实探测,对技能对进行移除、替换和重排,通过贝叶斯平滑聚合证据,发布经状态过滤的加权图用于任务条件扩展。该图离线构建,使用时无需改变下游智能体策略或任务接口。在ALFWorld ID-140和ScienceWorld U211上的6种LLM骨干模型测试中,CaSKG在模型与基准的12种组合中均取得最高任务得分;相较于Graph-of-Skills(GoS),其将6种模型在ScienceWorld上的宏平均得分从72.62提升至80.50,ALFWorld成功率从80.01%提升至86.79%,同时降低了两个基准上的平均环境步数。定性与消融分析进一步表明,校准后的边有助于检索保留前提条件、状态改变动作、验证例程及最终完成步骤。这些结果表明,边置信度校准是实现大规模紧凑可执行技能检索的有效途径。代码可访问:this https URL
英文摘要
Reusable skill libraries allow large language model (LLM) agents to reuse procedural knowledge across tasks, but they also turn memory access into a challenging retrieval problem. Full-library prompting preserves coverage at high context cost, vector retrieval returns compact neighborhoods but treats skills as independent text, and graph-based retrieval can recover workflow context only when the edges that carry relevance are reliable. We propose CaSKG, a counterfactual-causal skill graph framework that calibrates procedural relations before retrieval. CaSKG first builds a high-recall directed candidate graph from semantic, lexical, input/output, and structural evidence, with repair evidence and an optional LLM judge further refining candidate scores. It then applies direction-conditioned textual counterfactual probes that remove, substitute, and reorder skill pairs, aggregates the evidence with Bayesian smoothing, and publishes a state-filtered weighted graph for task-conditioned expansion. The graph is constructed offline and used without changing the downstream agent policy or task interface. Across six LLM backbones on ALFWorld ID-140 and ScienceWorld U211, CaSKG achieves the highest task score in all twelve combinations of model and benchmark. Relative to Graph-of-Skills (GoS), it improves the six-model macro-average ScienceWorld score from 72.62 to 80.50 and ALFWorld success from 80.01\% to 86.79\%, while reducing mean environment steps on both benchmarks. Qualitative and ablation analyses further show that calibrated edges help retrieval preserve prerequisites, state-changing actions, verification routines, and final completion steps. These results position edge-confidence calibration as an effective route to compact and executable skill retrieval at scale\footnote{Code is available at: https://github.com/ZhiyuanLi218/Caskg }.