arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24220cs.CV

文档检索感知分块(D-RAC):通过PDF归一化与多模态Markdown转换实现企业文档的通用检索感知摄入

Document Retrieval-Aware Chunking (D-RAC): Universal Retrieval-Aware Ingestion of Enterprise Documents via PDF Normalization and Multimodal Markdown Conversion

  • AI Research Team(AI研究团队)

机构由 AI 辅助整理,请以论文原文为准。

Uday Allu, Abhivanth Sivaprakash, Pratik Singh, Aman Manocha

AI总结:

D-RAC通过PDF归一化和多模态Markdown转换实现企业文档的通用检索感知摄入,在RAG基准上以零错误、低成本高效生成检索就绪块,显著优于代理式分块。

AI中文摘要:

检索增强生成(RAG)系统在企业知识库上必须摄入异构文档格式——PDF、Word文档、演示文稿和扫描件——这些内容被锁定在复杂的视觉布局、多栏页面和密集表格中。基于规则的提取和OCR会破坏阅读顺序、压平表格并丢失标题层级,而对提取文本进行完全代理式分块则会产生高昂的令牌成本和幻觉风险。我们提出了文档检索感知分块(D-RAC),这是我们网络检索感知分块(W-RAC)框架对任意文档格式的扩展。D-RAC首先将任何输入文档归一化为PDF,利用几乎每种格式都有忠实、确定性的PDF渲染这一事实。随后,单次多模态LLM传递将渲染页面转换为检索优化的Markdown——将表格改写为自包含的散文陈述并保留标题层级——之后分块完全按照W-RAC进行:确定性解析为ID可寻址单元,随后基于标识符而非文本进行轻量级基于LLM的分块规划。源文本在分块过程中从不重新生成,保留了W-RAC的成本、确定性和可观测性优势,同时使每种可渲染格式都成为一等输入。在RAG-Multi-Corpus基准的236个文档、795页PDF子集上(涵盖五个企业领域),D-RAC在72分钟内以零错误转换并分块整个语料库,生成1,748个检索就绪块。与使用前沿LLM的代理式分块相比,D-RAC将分块阶段的输出令牌减少95.7%,将分块成本降低77.8%(按GPT-4.1定价)至85.6%(按Gemini 2.5 Pro定价),并将分块时间减少75%。D-RAC可线性扩展到500页以上的文档。

英文摘要:

Retrieval-Augmented Generation (RAG) systems over enterprise knowledge bases must ingest heterogeneous document formats -- PDFs, Word documents, presentations, and scans -- whose content is locked inside complex visual layouts, multi-column pages, and dense tables. Rule-based extraction and OCR destroy reading order, flatten tables, and lose heading hierarchy, while fully agentic chunking over extracted text incurs high token costs and hallucination risk. We present Document Retrieval-Aware Chunking (D-RAC), an extension of our Web Retrieval-Aware Chunking (W-RAC) framework to arbitrary document formats. D-RAC first normalizes any input document into PDF, exploiting the fact that virtually every format has a faithful, deterministic PDF rendering. A single multimodal LLM pass then converts rendered pages into retrieval-optimized Markdown -- rewriting tables as self-contained prose statements and preserving heading hierarchy -- after which chunking proceeds exactly as in W-RAC: deterministic parsing into ID-addressable units followed by lightweight LLM-based chunk planning over identifiers rather than text. Source text is never regenerated during chunking, preserving W-RAC's cost, determinism, and observability benefits while unlocking every renderable format as a first-class input. On the 236-document, 795-page PDF subset of the RAG-Multi-Corpus benchmark spanning five enterprise domains, D-RAC converts and chunks the entire corpus in 72 minutes with zero errors, producing 1,748 retrieval-ready chunks. Compared to agentic chunking with frontier LLMs, D-RAC reduces chunking-stage output tokens by 95.7%, cutting chunking cost by 77.8% (GPT-4.1 pricing) to 85.6% (Gemini 2.5 Pro pricing) and chunking time by 75%. D-RAC scales linearly to documents of 500+ pages.

补充信息

↑