CHaystack:中文文档检索与视觉问答基准测试
CHaystack: Benchmarking Chinese Document Retrieval and VQA
浏览论文内容
中文总结 AI 辅助
本文介绍了CHaystack中文文档检索与视觉问答基准测试,涵盖四类文档。提出CDocRAG系统,用VLM相关性过滤器验证文档图像。评估开源模型发现Qwen系列在文本丰富文档表现佳,中文大规模DocumentVQA在文本编码方面有挑战,仍需改进。
中文摘要 AI 辅助
检索增强生成(RAG)在扩展大语言模型(LLMs)的记忆方面取得了显著进展,并且近期的进展进一步将RAG从纯文本设置推向多模态场景。在文档理解领域,文档视觉问答(DocumentVQA)已从单个文档的问答演变为大规模文档集合上的检索与生成管道。然而,专门针对中文大规模文档检索和问答的基准测试仍然缺乏。为了填补这一空白,我们引入了CHaystack,一个新的中文DocumentVQA基准测试,涵盖学术论文、广告、网页和真实世界拍摄的文档四类,能更全面地评估DocumentVQA系统。此外,我们提出了CDocRAG,一个中文DocumentVQA系统,它在答案生成前使用基于视觉语言模型(VLM)的相关性过滤器来验证检索到的文档图像。我们在CHaystack上评估了代表性的开源嵌入和生成模型。结果显示了各类别优势的明显对比:Qwen系列模型在网页和论文等文本丰富的文档上表现最佳,而其他模型仅在广告等视觉丰富的类别上取得有竞争力的结果,在文本密集的文档上则大幅下降。对于检索,Qwen3-VL的召回率达到71.91@1,而最佳的非Qwen模型仅为14.40。这些结果表明,CHaystack的核心挑战在于中文文本编码,并且中文大规模DocumentVQA仍有很大的改进空间。我们的代码和数据集可在该https网址获取。
英文摘要
Retrieval-augmented generation (RAG) has made substantial progress in extending the memory of large language models (LLMs), and recent advances have further pushed RAG from pure text settings toward multimodal scenarios. In the document understanding domain, document visual question answering (DocumentVQA) has evolved from question answering over a single document to retrieval-and-generation pipelines over large-scale document collections. However, a benchmark specifically designed for Chinese large-scale document retrieval and question answering is still lacking. To bridge this gap, we introduce CHaystack, a new Chinese DocumentVQA benchmark that covers four document categories, namely academic papers, advertisements, web pages, and real-world photographed documents, enabling a more comprehensive evaluation of DocumentVQA systems. In addition, we present CDocRAG, a Chinese DocumentVQA system that uses a VLM-based relevance filter to verify retrieved document images before answer generation. We evaluate representative open-source embedding and generation models on CHaystack. The results reveal a clear contrast in category-wise strengths: Qwen-family models perform best on text-rich documents such as webpages and papers, whereas other models only achieve competitive results on visually rich categories such as advertisements and degrade sharply on text-dense documents. For retrieval, Qwen3-VL reaches 71.91 Recall@1 while the best non-Qwen model achieves only 14.40. These results indicate that the core challenge of CHaystack lies in Chinese textual encoding, and that Chinese large-scale DocumentVQA still leaves substantial room for improvement. Our code and dateset is available at https://github.com/hanxi19/CHaystack.