EAR:面向检索增强生成开发的实体感知分区方法
EAR: Entity-Aware Partitioning Approach for Retrieval-Augmented Generation Development
浏览论文内容
中文总结 AI 辅助
EAR提出一种实体感知的语料库分区方法,通过锚点窗口检索减少检索词量,用于多项选择问答,虽准确率变化不显著,但提供了紧凑可检查的检索单元。
中文摘要 AI 辅助
检索增强生成(RAG)可以改善知识密集型问答,但第一个设计选择容易被忽视:源语料库应如何划分为可检索的单元?固定大小的分块常常返回长段落,这些段落与问题的关系仅是隐性的。我们提出了EAR,一种面向多项选择题解答(MCQA)的实体感知分区方法。EAR从问题、答案选项和语料库中提取归一化的表面锚点;检索匹配语料库锚点周围的局部窗口;并可通过抽取式摘要附加更大的父段落。我们在一个经过清洗的、由自动语料库支持启发式方法选出的153个问题的Massive Multitask Language Understanding(MMLU)风格子集上,并使用去污染后的公开教科书文本,对EAR进行了评估。在相同协议下,使用Mistral、Gemma和DeepSeek进行top-k = 3和top-k = 8的扫描,相对于分块,EAR实体窗口将检索到的单词减少了37.5-40.2%。观察到的准确率变化在top-k = 3时为+5.2、+1.3和-3.9个百分点,在top-k = 8时为+5.9、-3.3和-4.6个百分点;实体窗口的差异均无统计学显著性。本次贡献范围限定在方法论层面:EAR提供了一种紧凑且可检查的检索单元,而其基于规则的锚点提取器仍具有领域特定性,在迁移前需要单独验证。
英文摘要
Retrieval-augmented generation (RAG) can improve knowledge-intensive question answering, but the first design choice is easy to overlook: how should the source corpus be partitioned into retrievable units? Fixed-size chunks often return long passages whose relation to the question is only implicit. We introduce EAR, an Entity-Aware Partitioning approach for multiple-choice question answering (MCQA). EAR extracts normalized surface anchors from the question, answer options, and corpus; retrieves local windows around matching corpus anchors; and can attach a larger parent passage through an extractive summary. We evaluate EAR on a cleaned Massive Multitask Language Understanding (MMLU)-style subset of 153 questions selected by an automatic corpus-support heuristic and using decontaminated public textbook text. Across same-protocol top-k = 3 and top-k = 8 sweeps with Mistral, Gemma, and DeepSeek, EAR entity-window reduces retrieved words by 37.5-40.2% relative to chunks. Observed accuracy changes are +5.2, +1.3, and -3.9 points at top-k = 3, and +5.9, -3.3, and -4.6 points at top-k = 8; none of the entity-window differences is statistically significant. The scoped contribution is methodological: EAR provides a compact and inspectable retrieval unit, while its rule-based anchor extractor remains domain-specific and requires separate validation before transfer.
发表机构
- Georgia Institute of Technology(佐治亚理工学院)
- Ege University(爱琴大学)
机构由 AI 辅助整理,请以论文原文为准。