arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.04442cs.CLcs.AI

GRACE:面向专家参与的知识扩展的基于图的反思型智能体副驾驶引擎

GRACE: Graph-Grounded Reflective Agent Copilot Engine for Expert-in-the-Loop Knowledge Expansion

John Seon Keun Yi, Joshua R. Minot, Dokyun Lee

首次发表
浏览论文内容

中文总结 AI 辅助

GRACE是结合图结构表示与专家参与验证的框架,通过拆解LLM响应并分类断言,利用RoA机制分配资源,其知识库优于RAG基线,可在系统层面缓解LLM幻觉。

中文摘要 AI 辅助

部署在高风险场景中的大语言模型(LLM)经常生成看似合理但缺乏依据的断言。标准的检索增强生成(RAG)管道对此的改善有限,因为它们检索的是孤立的段落,未跟踪跨文档的证据关系或量化不确定性。我们提出GRACE(Graph-grounded Reflective Agent Copilot Engine,基于图的反思型智能体副驾驶引擎),该框架将LLM的响应拆解为原子断言,并针对加权二分图中的可信知识先验对这些断言进行依据核查。边权重编码每个断言与先验的接近程度,支持加权中心性分析,将断言分类为依据充分、被驳斥或处于边界状态。这种分类不仅能识别幻觉,还能识别模型知识前沿的新颖或有争议的断言。为了高效分配人类或智能体资源,我们制定了注意力回报率(RoA)目标,仅当断言的优先级加权不确定性超过验证成本时,才将其提交给专家审核。经专家验证的断言会被提升为新的证据锚点,形成验证器-LLM进化循环,该循环会在迭代过程中扩展知识库。我们在多个LLM以及涵盖通用和领域特定知识的数据集上对GRACE进行了评估。结果表明,我们的知识库作为检索的可靠基础,其性能优于RAG基线;且RoA框架能高效选择有价值的边界知识供专家验证。这些发现证明,基于图的表示与专家参与的验证相结合,可在系统层面而非生成层面缓解幻觉。代码可访问此httpsURL

英文摘要

Large language models deployed in high-stakes settings frequently generate plausible but ungrounded claims. Standard retrieval-augmented generation (RAG) pipelines offer limited remedy, since they retrieve isolated passages without tracking cross-document evidence relationships or quantifying uncertainty. We introduce GRACE (Graph-grounded Reflective Agent Copilot Engine), a framework that deconstructs LLM responses into atomic claims and grounds them against trusted knowledge priors within a weighted bipartite graph. Edge weights encode the closeness of each claim to the priors, enabling weighted centrality analysis that classifies claims as Grounded, Refuted, or Boundary. Such classification identifies not just hallucinations but also novel or contested claims at the frontier of the model's knowledge. To efficiently allocate human or agent resources, we formulate a Return on Attention (RoA) objective that defers a claim to expert review only when its priority-weighted uncertainty exceeds the cost of verification. Claims verified by experts are promoted to new evidence anchors, closing a validator-LLM evolutionary loop that expands the knowledge base across iterations. We evaluate GRACE across multiple language models and on datasets spanning both general and domain-specific knowledge. Our results show that our knowledge base serves as a reliable foundation for retrieval that outperforms RAG baselines, and that the RoA framework efficiently selects valuable boundary knowledge for expert verification. These findings demonstrate that graph-structured representations combined with expert-in-the-loop verification can mitigate hallucination at the system level rather than at the generation level. Code available at https://github.com/johnsk95/grace_code

发表机构

  • Boston University(波士顿大学)
  • MassMutual(恒康相互人寿保险公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑