超越片段化后再嵌入:信息检索中文档片段化策略的全面分类与评估
Beyond Chunk-Then-Embed: A Comprehensive Taxonomy and Evaluation of Document Chunking Strategies for Information Retrieval
AI总结:
本文研究了文档片段化策略在信息检索中的分类与评估,发现任务依赖性和分段方法差异影响效果,提出统一框架并提供公开基准。
AI中文摘要:
文档片段化是密集检索系统中的关键预处理步骤,但片段化策略的设计空间仍不明确。近期研究提出了几种并发方法,包括基于大语言模型的方法(例如DenseX和LumberChunker)和上下文化策略(例如晚期片段化),这些方法在分段前生成嵌入以保留上下文信息。然而,这些方法独立出现,并在基准测试中重叠最小,使得直接比较困难。本文重现已有的文档片段化研究,并提出一个系统框架,该框架沿两个关键维度统一现有策略:(1)分段方法,包括基于结构的方法(固定大小、基于句子和基于段落)以及语义引导和大语言模型引导的方法;以及(2)嵌入范式,这些范式决定了分段相对于嵌入的时间(预嵌入分段与上下文化分段)。我们的重现实验评估了这些方法在两个不同的检索设置中:文档内检索(针在 haystack 中)和文档内检索(标准信息检索任务)。我们的全面评估发现,最佳的片段化策略是任务依赖的:简单的基于结构的方法在文档内检索中优于基于大语言模型的替代方法,而LumberChunker在文档内检索中表现最佳。上下文化分段提高了文档内检索的有效性,但降低了文档内检索的效果。我们还发现,片段大小与文档内检索效果有中等相关性,但与文档内检索效果相关性较弱,这表明分段方法的差异并非单纯由片段大小驱动。我们的代码和评估基准已公开在(匿名)处。
英文摘要:
Document chunking is a critical preprocessing step in dense retrieval systems, yet the design space of chunking strategies remains poorly understood. Recent research has proposed several concurrent approaches, including LLM-guided methods (e.g., DenseX and LumberChunker) and contextualized strategies(e.g., Late Chunking), which generate embeddings before segmentation to preserve contextual information. However, these methods emerged independently and were evaluated on benchmarks with minimal overlap, making direct comparisons difficult. This paper reproduces prior studies in document chunking and presents a systematic framework that unifies existing strategies along two key dimensions: (1) segmentation methods, including structure-based methods (fixed-size, sentence-based, and paragraph-based) as well as semantically-informed and LLM-guided methods; and (2) embedding paradigms, which determine the timing of chunking relative to embedding (pre-embedding chunking vs. contextualized chunking). Our reproduction evaluates these approaches in two distinct retrieval settings established in previous work: in-document retrieval (needle-in-a-haystack) and in-corpus retrieval (the standard information retrieval task). Our comprehensive evaluation reveals that optimal chunking strategies are task-dependent: simple structure-based methods outperform LLM-guided alternatives for in-corpus retrieval, while LumberChunker performs best for in-document retrieval. Contextualized chunking improves in-corpus effectiveness but degrades in-document retrieval. We also find that chunk size correlates moderately with in-document but weakly with in-corpus effectiveness, suggesting segmentation method differences are not purely driven by chunk size. Our code and evaluation benchmarks are publicly available at (Anonymoused).