arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24092cs.AI

DocMIDE:在视觉丰富文档中学习多跳隐式推导

DocMIDE: Learning Multi-Hop Implicit Derivation in Visually Rich Documents

  • VX Real Limited(VX Real有限公司)
  • Hang Seng University of Hong Kong(香港恒生大学)
  • The Hong Kong Polytechnic University(香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Jeremy Cerwin Wang, Wai Kit Wong, Jeff Kai Tai Tang

AI总结:

DocMIDE通过微调框架,以规划-检索-推导结构和组相对策略优化奖励中间步骤,将视觉丰富文档中隐式字段提取准确率从70.8%提升至95.9%。

AI中文摘要:

现实世界的文档处理系统依赖于僵化、预定义的架构,然而关键目标字段往往在页面上缺乏直接的视觉对应物。提取这些隐式值需要多跳推导,例如聚合子类别或对视觉标记进行推理。现有方法虽然能处理显式文本片段或简单的隐式查询,但在标准微调后仍无法进行多跳视觉推理:模型要么检索到错误的视觉证据,要么正确检索到证据却跳过了推导的中间步骤。为解决这一问题,我们提出了DocMIDE,一个微调框架,用于训练紧凑的视觉语言模型在推导答案之前显式地检索视觉证据。DocMIDE将生成过程约束为“规划-检索-推导”结构,并使用组相对策略优化(Group Relative Policy Optimization)在四组件、基于规则的奖励下进行优化,该奖励根据经过验证的参考轨迹对输出格式、检索到的证据块、每个中间推导步骤以及最终值进行评分。在一个包含4,151对隐式提取的基准上,DocMIDE仅使用少量标注示例即将Qwen3.5-4B的准确率从70.8%提升至95.9%,并成功迁移到第二种骨干架构。仅使用监督演示无法在我们测试的任何预算下缩小这一差距;对中间步骤进行奖励才是关键。

英文摘要:

Real-world document processing systems rely on rigid, predefined schemas, yet critical target fields often lack direct visual counterparts on the page. Extracting these implicit values requires multi-hop derivation, such as aggregating sub-categories or reasoning over visual marks. While existing methods handle explicit text spans or simple implicit queries, they fail at multi-hop visual reasoning even after standard fine-tuning: models retrieve incorrect visual evidence, or retrieve it correctly and then skip the intermediate steps of the derivation. To address this, we introduce DocMIDE, a fine-tuning framework that trains compact vision-language models to retrieve visual evidence explicitly before deriving an answer. DocMIDE constrains generation to a plan-retrieve-derive structure and optimizes it with Group Relative Policy Optimization under a four-component, rule-based reward that scores output format, the retrieved evidence block, every intermediate derivation step, and the final value against a verified reference trace. On a 4,151-pair implicit extraction benchmark, DocMIDE raises accuracy from 70.8% to 95.9% on Qwen3.5-4B from only a small set of annotated examples, and transfers to a second backbone architecture. Supervised demonstrations alone do not close this gap at any budget we tested; rewarding the intermediate steps is what does.

↑