arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

NS-ST-GraphRAG:用于文学知识处理的神经符号时空图检索增强生成

NS-ST-GraphRAG: Neuro-Symbolic Spatio-Temporal GraphRAG for Literary Knowledge Processing

Zheng Kui Lin

arXiv 2609.05139首次发表:更新:

发表机构

Dalian Ocean University(大连海洋大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对长篇文学叙事的信息处理挑战,提出NS-ST-GraphRAG神经符号时空GraphRAG框架,构建首个中国古典文学多跳问答基准Red-Chamber-QA,实验显示其机械答案复现率与语义判断准确率优于基线模型。

AI 中文摘要

长篇文学叙事对检索增强生成(RAG)构成独特的信息处理挑战:相关证据分布在不同章节,关系随叙事时间演变,正确答案可能同时依赖时间、空间和关系约束。我们提出NS-ST-GraphRAG,这是一个神经符号时空GraphRAG框架,整合了本体引导提取、确定性约束检查、双时间坐标、空间场景属性及动态子图检索。该框架不从单一语料库级图中检索,而是选择查询时间和空间范围内有效的图状态,并将生成的答案基于可追溯证据。我们还引入Red-Chamber-QA,据我们所知,这是首个针对中国古典文学的开放多跳问答基准,包含时间、空间和通用问题类别、每部分证据跨度及确定性捷径控制。在120个问题的保留测试集上,NS-ST-GraphRAG的机械答案复现率为0.733,而冻结窗口基线为0.675,闭卷模型为0.083(McNemar精确检验p=0.092,方向有利但不显著);语义判断准确率为0.866,而冻结窗口基线为0.850。H2的预设约束类别条件未得到所做比较的支持。这些结果表明,时间图表示、约束提取和可审计评估如何整合为一个统一框架,用于长篇叙事的可验证知识处理。

英文摘要

Long-form literary narratives pose a distinctive information-processing challenge for retrieval-augmented generation: relevant evidence is distributed across chapters, relations evolve over narrative time, and correct answers may depend jointly on temporal, spatial, and relational constraints. We propose NS-ST-GraphRAG, a neuro-symbolic spatio-temporal GraphRAG framework that integrates ontology-guided extraction, deterministic constraint checking, dual temporal coordinates, spatial scene attributes, and dynamic sub-graph retrieval. Instead of retrieving from a single corpus-level graph, the framework selects the graph state valid for the temporal and spatial scope of a query and grounds generated answers in traceable evidence. We further introduce Red-Chamber-QA, to our knowledge the first open multi-hop question-answering benchmark for classical Chinese literature, with time-, space-, and general-question categories, per-part evidence spans, and deterministic shortcut controls. On a 120-question held-out split, NS-ST-GraphRAG achieves mechanical answer reproduction of 0.733 versus 0.675 for the frozen window baseline and 0.083 for a closed-book model (McNemar exact p = 0.092, directionally favorable but not significant); semantic-judge accuracy is 0.866 versus 0.850. The pre-specified constrained-category condition of H2 is not supported by the delivered comparison. These results show how temporal graph representation, constrained extraction, and auditable evaluation integrate into a unified framework for verifiable knowledge processing over long-form narrative.

CommentsSubmitted to Information Processing and Management

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑