复杂分块何时值得?大规模分块方法的多目标评估
When Is Complex Chunking Worth It? A Multi-Objective Evaluation of Chunking Methods at Scale
浏览论文内容
中文总结 AI 辅助
该研究针对长文档分块的多目标评估问题,在多语料库、多模型下对比八种分块策略,发现分块需兼顾检索性能与系统成本,应作为多目标设计决策。
中文摘要 AI 辅助
密集检索通常在将每个文档表示为单个嵌入的基准上进行评估,尽管现实中的检索系统常索引需要分块的长文档。在此类场景中,所选分块方法不仅影响检索质量,还影响索引吞吐量、查询延迟和内存使用。以往分块策略的比较主要聚焦于检索性能,对操作层面的权衡探索不足。为解决这些问题,我们在两个可扩展语料库、三个嵌入模型及多个语料库规模下评估八种代表性分块策略,同时测量检索有效性和系统级成本。结果显示,计算成本高的方法很少比简单分块提供持续增益,最佳策略取决于嵌入模型、数据集、语料库规模和目标检索指标;性能相近的方法操作成本可能差异显著,表明分块应视为多目标设计决策。
英文摘要
Dense retrieval is commonly evaluated on benchmarks that represent each document with a single embedding, even though real-world retrieval systems often index long documents that require chunking. In these settings, the chosen chunking method not only affects retrieval quality, but also indexing throughput, query latency, and memory usage. Prior comparisons of chunking strategies have mainly focused on retrieval performance, leaving operational trade-offs underexplored. To address these issues, we evaluate eight representative chunking strategies across two scalable corpora, three embedding models, and multiple corpus sizes, measuring both retrieval effectiveness and system-level costs. Our results show that computationally expensive methods rarely provide consistent gains over simpler chunking. Instead, the best performing strategy depends on the embedding model, dataset, corpus size, and target retrieval metric. Methods with similar performance can also differ substantially in operational cost, showing that chunking should be seen as a multi-objective design decision.