arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

结合标题链前缀的结构感知语义分块:1600次查询评估及文本变换消融实验中的测量陷阱

Structure-Aware Semantic Chunking with Title-Chain Prefixes: A 1600-Query Evaluation and the Measurement Trap in Text-Transform Ablations

Yang Yang

arXiv 2608.00824首次发表:更新:

AI 中文总结

本文提出仅在分块侧的三阶段语义分块流水线,经1600次查询评估提升RAG的MRR@5,还发现分块研究中存在的测量陷阱并推荐检索时每个候选前缀评估的方向。

AI 中文摘要

分块是检索增强生成(RAG)的第一步,也是最关键的一步:所有下游检索决策都继承分块边界。我们提出一种仅在分块侧的三阶段语义分块流水线——标题拆分、语义合并与标题链前缀,该流水线无需额外的大语言模型(LLM)调用:标题链复用文档自身的标题层级,而非生成摘要。在对生产环境的Markdown知识库开展的1600次查询分层评估中,该流水线使MRR@5在全部查询集上从0.374提升至0.463(+23.8%),在可回答子集(n=563)上从0.828提升至0.925(+11.7%),双标注者的Cohen's kappa为0.45(未加权,16000个评分对)。随后我们报告了尝试过的方案与失败的情况:三个查询侧或架构层面的后续方案是设计死胡同(未评估——无可比运行产物),一个是测量失败(前缀权重衰减),还有一个是协议层面的失败,这是本文核心的方法学发现。在同一池的前缀开启/关闭消融实验中,双标注者一致性在相同提示下从kappa 0.45骤降至0.04——剥离标题链上下文会剥离标注者就相关性达成一致所需的消歧信号。这种测量陷阱使分块研究中的一种常见评估实践失效,并推动了检索时每个候选前缀评估,这是我们从所有证据中推荐的方向。

英文摘要

Chunking is the first and most consequential step in retrieval-augmented generation (RAG): every downstream retrieval decision inherits the chunk boundaries. We present a three-stage, chunk-side-only semantic chunking pipeline---header-split, semantic merge, and title-chain prefixing---that costs zero additional LLM calls: the title chain reuses the document's own header hierarchy instead of a generated summary. On a 1600-query stratified evaluation over a production Markdown knowledge base, the pipeline improves MRR@5 from 0.374 to 0.463 (+23.8%) on the full set and from 0.828 to 0.925 (+11.7%) on the answerable subset (n=563), with dual-annotator Cohen's kappa 0.45 (unweighted, 16,000 score pairs). We then report what we tried and what failed: three query-side or architecture-level follow-ups are design dead ends (unevaluated---no comparable run artifacts), one measured failure (prefix weight decay), and one protocol-level failure that is the paper's central methodological finding. In a same-pool prefix on/off ablation, dual-annotator agreement collapsed from kappa 0.45 to 0.04 under identical prompts---stripping the title-chain context strips the disambiguation signal annotators need to agree on relevance. This measurement trap invalidates a common evaluation practice in chunking research and motivates retrieval-time per-candidate prefix evaluation, the direction we recommend from all our evidence.

Comments8 pages, 1 table. Replication package: DOI 10.5281/zenodo.21744653

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑