通过检索增强大语言模型利用外部知识进行历史文献修复
Leveraging External Knowledge for Historical Document Restoration via Retrieval-Augmented Large Language Models
浏览论文内容
中文总结 AI 辅助
针对历史文献因物理损坏难以辨认及现有修复方法在恢复命名实体上的局限,提出利用检索增强大语言模型的修复框架,结合预训练模型隐含知识与外部上下文,实验表明该方法性能优于基线,是历史记录分析实用工具。
中文摘要 AI 辅助
历史文献是宝贵的知识档案,但常因物理退化和损坏而难以辨认。现有基于掩码语言建模的修复方法虽能有效利用局部上下文,但在恢复需要外部历史知识的命名实体时存在困难。为解决这一局限,我们引入了一种利用检索增强生成(RAG)大语言模型的历史文献修复新框架。通过将预训练大语言模型的隐含知识与明确检索到的外部上下文相结合,我们的模型ARI有效缓解了推断上下文相关专有名词的挑战。在韩国历史文献上的大量实验表明,我们的方法显著优于基线,在恢复普通字符和命名实体方面都有大幅提升。此外,包括专家评估在内的综合评估证实,ARI是领域专家的实用工具,有望加速历史记录分析。
英文摘要
Historical documents act as invaluable knowledge archives but often suffer from illegibility due to physical deterioration and damage. While existing restoration methods based on masked language modeling effectively utilize local context, they struggle to restore named entities that require external historical knowledge. To address this limitation, we introduce a novel framework for historical document restoration that leverages large language models with retrieval-augmented generation (RAG). By combining the implicit knowledge of pre-trained LLMs with explicitly retrieved external context, our model ARI effectively mitigates the challenge of inferring context-dependent proper nouns. Extensive experiments on Korean historical documents demonstrate that our approach significantly outperforms baselines, achieving substantial gains in restoring both general characters and named entities. Furthermore, comprehensive evaluations including expert assessments confirm that ARI serves as a practical tool for domain experts, promising to accelerate the analysis of historical records.
发表机构
- Kangwon National University(江原国立大学)
机构由 AI 辅助整理,请以论文原文为准。