arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GraphQAG:一种用于问答对生成的知识图谱引导型可视分析框架

GraphQAG: A Knowledge-Graph-Guided Visual Analytics Framework for Question-Answer Pairs Generation

Yize Li, Ruiqi Yu, Tianya Pan, Ningxin Li, Songyue Li, Xiangyang Wu, Jinchang Li, Zhiguang Zhou

arXiv 2607.27182首次发表:更新:

AI 中文总结

该研究针对长文档知识碎片化的问题,提出GraphQAG知识图谱引导型可视分析框架,经评估可帮助用户优化问答对集合,提升问答对的全面性与可信度。

AI 中文摘要

问答(QA)对广泛应用于知识库构建、问答系统以及大语言模型(LLM)的后训练中。然而,长文档中的重要知识往往分布在多个段落中,并通过复杂的实体关系相连接。这种碎片化且具有关联性的知识对现有问答对生成方法构成了重大挑战,这些方法往往无法充分覆盖文档核心内容、跨段落语义关联以及多实体关系。我们提出GraphQAG,这是一种用于从长文档生成高质量问答对的知识图谱引导型可视分析框架。GraphQAG遵循三阶段工作流程:首先,它将文档分割为段落,提取显著实体与关系,以此构建文档知识图谱;其次,它基于实体、关系和多跳路径构建图式生成空间,以约束并引导基于LLM的问答对生成;最后,它将知识图谱作为交互式可视化表示,使用户能够探索文档知识结构、检查生成问答对的覆盖范围与证据来源,并通过基于图的交互迭代优化问答对集合。我们通过包含16名参与者的用户研究、两个案例研究以及专家访谈对GraphQAG进行评估,结果表明GraphQAG能有效支持用户识别知识覆盖缺口、检查生成的问答对并优化问答对集合。这些发现证明,将知识图谱、基于LLM的生成以及可视分析相结合,对于从长文档生成更全面、可信的问答对具有实用价值。

英文摘要

Question-answer (QA) pairs are widely used in knowledge base construction, question-answering systems, and the post-training of large language models (LLMs). However, important knowledge in long documents is often distributed across multiple paragraphs and connected through complex entity relationships. Such fragmented and relational knowledge poses substantial challenges for existing QA generation methods, which often fail to adequately cover core document content, cross-paragraph semantic connections, and multi-entity relationships. We present GraphQAG, a knowledge graph-guided visual analytics framework for generating high-quality QA pairs from long documents. GraphQAG follows a three-stage workflow. First, it constructs a document knowledge graph by segmenting the document into paragraphs and extracting salient entities and relations. Second, it builds a graph-based generation space from entities, relations, and multi-hop paths to constrain and guide LLM-based QA generation. Third, it uses the knowledge graph as an interactive visual representation, enabling users to explore document knowledge structures, inspect the coverage and evidence provenance of generated QA pairs, and iteratively refine the QA pair set through graph-based interactions. We evaluated GraphQAG through a user study with 16 participants, two case studies, and expert interviews. The results indicate that GraphQAG effectively supports users in identifying knowledge coverage gaps, examining generated QA pairs, and refining the QA pair set. These findings demonstrate the usefulness of combining knowledge graphs, LLM-based generation, and visual analytics for producing more comprehensive and trustworthy QA pairs from long documents.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑