发表机构
Zhejiang University; Nanyang Technological University; Hangzhou City University(浙江大学; 南洋理工大学; 杭州城市大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
GraphCert提出一种自认证方法,通过生成并认证证据准则,在训练后阶段引导图智能体推理,无需人工标注或外部LLM,且在多个图推理领域上优于更大模型。
AI 中文摘要
图智能体通过多步与图工具的交互,扩展了大型语言模型(LLM)主动探索和推理知识图谱的能力。然而,训练有能力的图智能体通常需要大量的问答对和推理轨迹,这些数据的人工构建成本高昂且难以扩展。此外,使用专有LLM生成此类监督数据还可能将敏感图数据暴露给外部服务。因此,我们提出GraphCert,在训练后阶段通过认证证据准则引导智能体图推理。具体而言,由生成控制引导的引导式图测验器生成基于图的问答对并标记支持证据,这些证据经过执行认证和语义筛选。被接受的证据随后被规范化成认证证据准则,在GRPO训练期间,这些准则用于奖励图求解器的证据对齐和答案正确性。在GRBENCH的五个图推理领域上的实验表明,GraphCert始终优于规模大得多的LLM智能体和训练后方法。此外,我们的分析表明,学习到的策略在不同图领域间具有稳健的迁移性,这表明GraphCert获得了可复用的图推理能力,而非特定领域的模式。这些结果确立了可执行的自认证作为自训练紧凑图推理智能体的有效方法。我们的代码将公开发布。
英文摘要
Graph agents extend large language models (LLMs) with the ability to actively explore and reason over knowledge graphs through multi-step interactions with graph tools. However, training capable graph agents typically requires large collections of question-answer pairs and reasoning trajectories, whose manual construction is costly and difficult to scale. Moreover, employing proprietary LLMs to generate such supervision further risks exposing sensitive graph data to external services. Therefore, we propose GraphCert to bootstrap agentic graph reasoning with certified evidence rubrics during post-training. Specifically, the Bootstrapped Graph Quizzer guided by generation controls produces graph-grounded QA pairs and marks supporting evidence, which undergo execution certification and semantic curation. The accepted evidence is then canonicalized into certified evidence rubrics that later reward Graph Solver evidence alignment alongside answer correctness during GRPO training. Experiments on five graph reasoning domains in GRBENCH demonstrate that GraphCert consistently outperforms substantially larger LLM agents and post-training method. Furthermore, our analysis demonstrates that the learned policy transfers robustly across heterogeneous graph domains, suggesting that GraphCert acquires reusable graph-reasoning capabilities rather than domain-specific patterns. These results establish executable self-certification as an effective approach to self-training compact graph reasoning agents. Our code will be made publicly available.
CommentsUnder review