arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31660cs.CLcs.AIcs.IR

解析器、分块与嵌入在印度政府监管文档检索增强生成中的交互作用

Parser, Chunking, and Embedding Interactions in Retrieval-Augmented Generation over Indian Government Regulatory Documents

Shubham Kumar Singh

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过因子实验联合评估解析器、分块和嵌入模型在印度政府监管文档RAG中的交互作用,发现无单一检索器占优且证据保留近饱和,检索差异主要由排序质量决定。

中文摘要 AI 辅助

检索增强生成(RAG)流水线通常由独立选择的组件组装而成——文档解析器、分块策略和嵌入模型——然而这些选择很少被联合评估,且确实结合这些组件的评估通常仅在单一文档或语料库上进行。我们针对3种解析器、3种分块策略和5种稠密嵌入模型,连同稀疏BM25基线,进行了一项受控因子研究,在800个问题实例上进行了评估,每个实例有一个或多个必需的证据字符串,证据字符串自动对照源文本验证,并对10%的随机样本进行人工审查,覆盖了四份结构不同的印度中央政府监管文件。我们对结果集(72,000行)拟合了具有文档-查询级随机截距的线性混合效应模型,在匹配的稠密和稀疏配置之间进行了Holm校正的配对比较,并为所有54种独特的检索器配置报告了聚类自助置信区间。我们发现没有单一的检索器家族在所有文档中占主导地位;解析器和分块器的选择存在显著交互;MPNet-base持续表现不佳,在表格衍生问题上存在严重故障模式;语料库显示出接近饱和的证据保留上限,超过98%,表明检索差异主要由排序质量驱动,而非摄取过程中的信息损失。我们还报告了嵌入维度和块大小/重叠的消融实验以及效率/质量帕累托分析。我们发布了完整的评估框架、语料库清单和800问题基准。

英文摘要

Retrieval-augmented generation (RAG) pipelines are typically assembled from independently-chosen components -- a document parser, a chunking strategy, and an embedding model -- yet these choices are rarely evaluated jointly, and evaluations that do combine them are usually run on a single document or corpus. We present a controlled factorial study of 3 parsers, 3 chunking strategies, and 5 dense embedding models, together with a sparse BM25 baseline, evaluated against 800 question instances, each with one or more required evidence strings, with evidence strings automatically validated against source text and a 10% random sample manually reviewed, across four structurally distinct Indian central-government regulatory documents. We fit linear mixed-effects models with document-query-level random intercepts to the resulting 72,000-row result set, run Holm-corrected paired comparisons between matched dense and sparse configurations, and report clustered bootstrap confidence intervals for all 54 unique retriever configurations. We find that no single retriever family dominates across documents; parser and chunker choice interact significantly; MPNet-base is a consistent underperformer with a severe failure mode on table-derived questions; and the corpus exhibits a near-saturated evidence-preservation ceiling above 98%, indicating that retrieval differences are driven primarily by ranking quality rather than information loss during ingestion. We additionally report embedding-dimension and chunk-size/overlap ablations and an efficiency/quality Pareto analysis. We release our full evaluation harness, corpus manifest, and 800-question benchmark.

补充信息

↑