arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.10170cs.IRcs.AI

文档结构是否有助于稠密检索?两个语料库上四种机制的安慰剂对照消融研究

Does Document Structure Help Dense Retrieval? A Placebo-Controlled Ablation of Four Mechanisms Across Two Corpora

Andrey Kuehlkamp, Priscila Correa Saboia Moreira, Samuel Rund

首次发表
浏览论文内容

中文总结 AI 辅助

通过安慰剂对照消融实验,验证文档结构处理对稠密检索的贡献,发现结构对齐分块和真实标题路径有效,而朴素分层检索有害,效应虽小但跨语料库稳健。

中文摘要 AI 辅助

检索增强生成系统越来越依赖文档结构处理:结构对齐的分块、LLM生成的块上下文、标题路径元数据以及分层两阶段检索。已有研究在不同语料库、嵌入器和指标上分别支持每种方法,但均未控制一个共同的混杂因素:任何添加到块前面的文本都会扰动其嵌入。我们提出了一种机制隔离的消融研究,在统一协议下测试所有四种处理,匹配各条件下的块大小,并添加语义上无效的安慰剂——结构有效但在文档间打乱的标题路径。我们使用覆盖率感知的nDCG进行检索评分,并通过文档聚类自助法与Holm校正测试四个预注册对比,在两个相距较远的语料库上进行:200篇维基百科特色文章(951个查询)和1,585篇QASPER论文(4,303个问题)。组织性有帮助,原因是内容而非词元:具有真实标题路径的结构对齐块优于带上下文的固定窗口(+0.022 / +0.012 cov-nDCG@10)和安慰剂(+0.010 / +0.016)。朴素的两阶段分层检索效果较差(-0.033 / -0.015),可追溯到第一阶段章节召回率。在维基百科上,真实结构优于LLM生成的结构,但在QASPER上并非如此。效应量较小($dz$ 0.06-0.11),但经Holm校正后显著且跨语料库一致。

英文摘要

Retrieval-augmented generation systems increasingly rely on document-structure treatments: structure-aligned chunking, LLM-generated chunk contexts, heading-path metadata, and hierarchical two-stage retrieval. Separate studies support each on different corpora, embedders, and metrics, and none control for a shared confound: any text prepended to a chunk perturbs its embedding. We present a mechanism-isolating ablation testing all four treatments under one protocol, matching chunk sizes across conditions and adding a semantically null placebo---heading paths that are structurally valid but shuffled across documents. We score retrieval with a coverage-aware nDCG and test four pre-registered contrasts via document-clustered bootstrap with Holm correction, on two distant corpora: 200 Wikipedia Featured Articles (951 queries) and 1,585 QASPER papers (4,303 questions). Organization helps, and the cause is content, not tokens: structure-aligned chunks with real heading paths beat contextualized fixed windows (+0.022 / +0.012 cov-nDCG@10) and the placebo (+0.010 / +0.016). Naive two-stage hierarchical retrieval hurts (-0.033 / -0.015), traceable to first-stage section recall. Gold structure beats LLM-induced structure on Wikipedia but not on QASPER. Effects are small ($dz$ 0.06-0.11) but Holm-significant and consistent across corpora.

发表机构

  • Center for Research Computing, University of Notre Dame(圣母大学研究计算中心)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑