arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

土耳其检索增强生成(RAG)系统的分块与嵌入策略比较

Comparing Chunking and Embedding Strategies for Turkish RAG Systems

Mustafa Sertaç Türkel, Fatma Nur Korkmaz, Ahmet Tuğrul Bayrak

arXiv 2608.26192首次发表:更新:

发表机构

Data Science and Innovation Ata Technology Platforms(阿塔技术平台数据科学与创新部门)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对土耳其语RAG系统,对比三种分块策略、五种嵌入模型和两种生成器LLM,发现布局感知分块可压缩嵌入模型差异,最优配置依内容类型而定,最佳单个组件无法组合出最优整体配置,准确率达87.0%。

AI 中文摘要

文档被分割为可检索分块的方式以及这些分块的嵌入方式会强烈影响检索增强生成(RAG)的质量,但针对土耳其等形态丰富的语言,尚未对这两种方式进行系统研究。我们针对土耳其文档问答任务,在三种具有不同布局的文档上,比较了三种分块策略(固定长度、语义、布局感知Docling)、五种嵌入模型和两种生成器大语言模型(LLM)的表现。完全交叉设计产生了9000个分级问答评估结果,每个结果均由独立的评判模型打分,组件比较通过经Holm校正的配对McNemar检验进行测试。研究得出四项发现:分块策略决定了嵌入选择的重要程度,布局感知分块将现代嵌入模型之间的差异压缩至约1个百分点;三种领先的嵌入模型在统计上无差异,因此语言专业化未产生可测量的检索优势;速度更快的生成器并非更准确的那个;且最优配置取决于内容类型,因为布局感知分块对包含表格的文档的帮助远大于对散文的帮助。因此,最优的单个组件无法组合成最优的完整配置,该完整配置的准确率达到87.0%。

英文摘要

Retrieval-Augmented Generation conditions a language model on chunks retrieved from a document collection. Its accuracy is therefore limited by the chunking and embedding stages that determine what can be retrieved. We compare Turkish document question answering across three chunking strategies (fixed-length, semantic, and layout-aware Docling), five embedding models, and two LLMs, over three documents with contrasting layouts. Every configuration answers the same question set, which allows component effects to be separated by paired testing rather than inferred from separate benchmarks. The fully crossed design yields 9{,}000 graded question-answer evaluations, each scored by an independent judge model, and component comparisons are tested by paired McNemar tests under Holm correction. The three leading embedding models are statistically indistinguishable, so language specialization yields no measurable retrieval advantage. The faster LLM is not the more accurate one. The preferred configuration depends on content type, since layout-aware chunking helps table-heavy documents far more than text-heavy ones.

CommentsAccepted to INTCEC 2026. This is the author's pre-print version. The final authenticated version will be available through the conference proceedings

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑