EnSI-RAG:面向长文档问答的实体结构索引增强检索生成框架
EnSI-RAG: Entity-Structure-Indexed Retrieval-Augmented Generation for Long-Document Question Answering
AI总结:
该研究针对长文档问答的RAG方法缺陷,提出EnSI-RAG框架,构建以实体为中心的索引,在Loong和Oolong数据集上平均准确率达78.24,较基准提升6.62个百分点,验证了其有效性。
AI中文摘要:
针对长文档的问答(QA)任务仍具挑战性,因为相关证据可能跨越多个实体及其关系。现有增强检索生成(RAG)方法通常将文档索引为原始文本块,并通过嵌入相似度进行检索;当文本块边界将实体与支撑证据分隔开,或问题需要在整个语料库中进行多跳推理时,其性能会下降。我们提出EnSI-RAG(Entity-Structure-Indexed Retrieval-Augmented Generation,实体结构索引增强检索生成),这是一种构建与查询无关、以实体为中心的索引的框架。每个记录(e, t, k, v)表示一个实体e、其类型t、{属性、关系、方面}中的语义类别k以及值v,同时保留与原始源段落的链接。在查询阶段,这些记录作为检索句柄,大型语言模型(LLM)将检索到的段落合成为最终答案。该设计将证据定位与答案合成分离,同时保留可追溯的源证据。在Loong和Oolong数据集上,EnSI-RAG的平均准确率为78.24;相对于作为参考的已发布基准分数,其高出6.62个百分点,表明其在这些设置下的有效性。代码可从指定URL获取。
英文摘要:
Question answering (QA) over long, connected documents remains challenging because relevant evidence may span multiple entities and their relationships. Existing retrieval-augmented generation (RAG) methods typically index documents as raw chunks and retrieve them through embedding similarity. Their performance degrades when chunk boundaries separate entities from supporting evidence or when a question requires multi-hop reasoning across the corpus. We propose EnSI-RAG (Entity-Structure-Indexed Retrieval-Augmented Generation), a framework that constructs a query-independent, entity-centered index. Each record (e, t, k, v) represents an entity e, its type t, a semantic category k in {property, relation, aspect}, and a value v, while retaining links to the original source passages. At query time, these records serve as retrieval handles, and an LLM synthesizes the retrieved passages into the final answer. This design separates evidence localization from answer synthesis while preserving traceable source evidence. Across Loong and Oolong, EnSI-RAG achieves an average accuracy of 78.24. Relative to the published baseline scores used as references, this is 6.62 points higher, suggesting its effectiveness across these settings. The code is available at https://github.com/RamonMeng/EnSI-RAG.