arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.01852cs.IRcs.AIcs.CL

评估学术文本中检索增强生成的文本分块策略

Evaluating Chunking Strategies for Retrieval-Augmented Generation on Academic Texts

Valentin J. J. Kreileder, Johannes Reisinger, Andreas Fischer

AI总结:

研究在检索增强生成系统中,基于聚类的语义分块是否优于固定大小和递归分块,通过RAGAs框架在长结构化学术论文上评估检索和答案质量,发现聚类分块未超越简单策略。

AI中文摘要:

检索增强生成(RAG)系统利用大型语言模型(LLMs)的问答能力访问其参数之外的信息。我们使用检索增强生成评估(RAGAs)框架,在长结构化学术论文上评估基于聚类的语义分块是否比固定大小和递归分块提高检索和答案质量。基于RAGAs的忠实度在此设置中显示出有限的可靠性。固定与文档特定问题的表现差异很大,可能与文档格式和预处理有关。在测试配置下,基于聚类的分块并未优于更简单的策略。

英文摘要:

Retrieval-Augmented Generation (RAG) systems use the question-answering capabilities of Large Language Models (LLMs) to access information outside their parameters. We evaluate if cluster-based semantic chunking improves retrieval and answer quality compared to fixed-size and recursive chunking evaluating on long, structured academic theses using the Retrieval Augmented Generation Assessment (RAGAs) framework. RAGAs based faithfulness shows limited reliability in this setup. Performance on fixed versus document specific questions varied substantially, likely related to the formatting of documents and preprocessing. Under the tested configuration, cluster-based chunking did not outperform simpler strategies.

↑