arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DeCoRAG:用于复杂文档理解的认知解耦和语义感知裁剪

DeCoRAG: Cognitive Decoupling and Semantic-Aware Cropping for Complex Document Understanding

Shuo Wang, Kai Zhang, Wenyuan Huang, Yizheng Yu, Xia Liao, Junming Su, Qing Wang, Fang Xi

arXiv 2607.24554首次发表:更新:

发表机构

QiYuanLab; Beijing University of Posts and Telecommunications(启元实验室; 北京邮电大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对复杂文档理解中多模态检索增强生成的难题,提出DeCoRAG方法,通过认知解耦、建立语义锚及区域感知裁剪机制,提升语义通过率,减少提示令牌,在复杂文档基准测试中取得良好效果。

AI 中文摘要

推进用于复杂文档理解的多模态检索增强生成(RAG)面临准确性和效率的双重困境,尤其是在图形RAG中。处理结构稀疏但视觉密集的布局会导致计算成本过高和灾难性幻觉。多模态图形RAG管道依赖于假设视觉语言模型(VLM)能解决高密度布局中稀疏语义的图形构建阶段,但这一假设存在问题。我们引入DeCoRAG,它将知识处理从耦合的视觉语义推理转变为“认知解耦”。其图形构建阶段建立宏观语义锚以中和注意力下沉,驱动区域感知修剪和裁剪机制,提高了语义通过率,减少了离线图形构建提示令牌。

英文摘要

Advancing multimodal retrieval-augmented generation (RAG) for complex document understanding presents a formidable dual dilemma of accuracy and efficiency, particularly in graph RAG. Processing structurally sparse yet visually dense layouts, such as extracting a tiny data marker from a financial chart, often incurs computationally prohibitive token overhead while still triggering catastrophic hallucination. However, multimodal Graph RAG pipelines rely on graph-construction stages that assume Vision-Language Models (VLMs) can resolve sparse semantics within high-density layouts. We challenge this assumption, revealing that forcing VLMs to localize visual evidence, interpret semantics, and extract relations triggers a "Visual Attention Sink," a mechanism driving catastrophic semantic loss, while full-page processing incurs massive computational overhead. Controlled interventions verify that this failure is boundary-driven rather than content-specific and that semantic anchoring mitigates it. To fundamentally correct this flawed paradigm, we introduce DeCoRAG, a multimodal Graph RAG pipeline that shifts knowledge processing from coupled visual-semantic reasoning to "Cognitive Decoupling." Rather than passively processing raw pixels, its graph-construction stage establishes a macroscopic Semantic Anchor to neutralize the attention sink. This anchor subsequently drives our Region-Aware Pruning and Cropping (RAP-Crop) mechanism, shifting the reasoning space from dense, noisy backgrounds to purified, intent-driven semantic clusters. The resulting graph supports hybrid retrieval and answer generation. Across complex document benchmarks, DeCoRAG improves the semantic pass rate by up to 12.5 percentage points over the strongest baseline and generalizes to DocVQA. RAP-Crop reduces offline graph-construction prompt tokens by 40.8% without sacrificing end-to-end accuracy.

Comments11 pages, 4 figures, 8 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑