arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过查询推理和执行对知识图谱中的历史文档进行鲁棒解释

Robust Interpretation of Historical Documents in Knowledge Graphs Through Query Inference and Execution

Sebastià Nicolau, Adrià Molina, Oriol Ramos Terrades, Josep Lladós

arXiv 2607.24475首次发表:更新:

AI 中文总结

研究如何在利用大语言模型泛化能力时保证可靠性,提出智能检索系统,比较传统RAG与智能GraphRAG架构,引入半符号框架,通过单词识别与代码生成协作构建鲁棒检索查询,提升历史文档分析准确性与可验证性。

AI 中文摘要

大语言模型(LLMs)的出现重新定义了用户在数字环境中与信息交互的方式。但其广泛且不加区分的整合引发了对可靠性和可信度的担忧,在访问数字图书馆和历史档案时尤为关键。本文提出一个智能检索系统,旨在在保持与无约束LLMs相关的灵活性的同时,更准确、可验证地访问历史数据。作为对历史文档分析的贡献,比较了传统检索增强生成(RAG)和智能GraphRAG架构在现实条件下提供历史信息的能力。引入了一个半符号框架,将用于OCR后校正的单词识别技术与知识图谱表示相结合。单词识别和代码生成之间的交错协作使智能体能够构建对错误解释和幻觉具有鲁棒性的强大检索查询,同时在历史文档分析中常见的噪声和不确定性阻碍精确检索时仍利用近似搜索。

英文摘要

The emergence of Large Language Models (LLMs) has redefined how users interact with information in digital environments. However, their widespread and often indiscriminate integration has raised significant concerns regarding reliability and trustworthiness issues that are particularly critical when accessing digital libraries and historical archives. How can one leverage the generalization capacity of an LLM without losing the level of accountability required for an archival institution? In this paper, we present an agentic retrieval system designed to deliver more accurate and verifiable access to historical data while preserving much of the flexibility associated with unconstrained LLMs. As a contribution to historical document analysis, we compare traditional Retrieval-Augmented Generation (RAG) with an agentic GraphRAG architecture in their ability to deliver historical information under realistic conditions, including the presence of OCR and transcription errors. We introduce a semi-symbolic framework that integrates word-spotting techniques for post-OCR correction with a knowledge graph representation that enables the agent to access information through synthesized queries. The interleaved collaboration between word spotting and code generation allows the agent to construct strong retrieval queries that are robust to misinterpretation and hallucination, while still leveraging approximate search when noise and uncertainty, common in historical document analysis, would otherwise hinder precise retrieval.

CommentsAccepted at ICDAR2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑