DocMemo:面向多模态文档理解的概率记忆引导检索动态证据发现框架
DocMemo: Dynamic Evidence Discovery via Probabilistic Memory-Guided Retrieval for Multi-Modal Document Understanding
浏览论文内容
中文总结 AI 辅助
针对长文档理解的静态检索与跨轮记忆局限,提出记忆引导框架DocMemo,通过三层记忆与多轮优化实现动态证据发现,在3个基准上取得SOTA性能。
中文摘要 AI 辅助
长文档理解需要在数百页中定位稀疏且异构的证据,但现有系统受限于静态检索和脆弱的跨轮次记忆。主流单轮方法从一开始就固定了前k页集合,难以从早期检索错误中恢复;近期迭代方法允许多轮证据获取,但未研究跨轮次状态的传播机制,难以跟踪页面相关性的动态变化。为解决这些局限,我们提出DocMemo,一种记忆引导框架,将长文档推理建模为动态证据探索。DocMemo维护三层检索状态,包括文档模式记忆、页面信念记忆和问题情景记忆,分别捕获结构先验、动态相关性估计和查询特定推理轨迹。推理过程中,DocMemo通过带Thompson采样的贝叶斯页面信念更新、空间邻近传播和结构感知自适应粒度证据访问,持续优化跨轮次页面选择,同时用细粒度视觉区域补充页面级证据。在3个基准上的实验表明,DocMemo达到了当前最优性能,验证了结构化记忆和动态页面信念更新的有效性。代码可在指定URL获取。
英文摘要
Long-document understanding requires locating sparse and heterogeneous evidence across hundreds of pages, yet existing systems remain limited by static retrieval and fragile cross-round memory. Mainstream single-round methods commit to a fixed top-$k$ page set at the outset and struggle to recover from early retrieval errors; recent iterative approaches allow multi-round evidence acquisition, but they do not investigate the propagation mechanism of cross-round states, making it difficult to track the dynamic changes in page relevance. To address these limitations, we propose DocMemo, a memory-guided framework that formulates long-document reasoning as dynamic evidence exploration. DocMemo maintains a tri-level retrieval state consisting of Document Schema Memory, Page Belief Memory, and Question Episodic Memory, which respectively capture structural priors, dynamic relevance estimation, and query-specific reasoning trajectories. During reasoning, DocMemo continuously refines cross-round page selection through Bayesian page belief updating with Thompson sampling, spatial proximity propagation, and structure-aware adaptive-granularity evidence access, while supplementing page-level evidence with fine-grained visual regions. Experiments on 3 benchmarks show that DocMemo achieves state-of-the-art performance and validate the efficacy of structured memory and dynamic page belief updating. Code is available at https://github.com/Harrygof/DocMemo.
发表机构
- Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
- Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)
机构由 AI 辅助整理,请以论文原文为准。