发表机构
Georgia Tech; CMU; Adobe; ZJU; NUIST; SEU; Physion Labs; Columbia University(佐治亚理工学院; 卡内基梅隆大学; 奥多比公司; 浙江大学; 南京信息工程大学; 东南大学; Physion实验室; 哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出面向推理的ADOPD 2026数据集,将多类文档元素转化为视觉锚点,支持区域语义标注、统一视觉-语言接地等能力,解决现有模型在文档计数任务上的不足,推动文档理解向锚点接地的智能方向发展。
AI 中文摘要
现有的文档理解基准大多聚焦于页面元素的定位,但现实世界的文档智能需要模型联合推理区域语义、空间关系和视觉结构。我们提出ADOPD 2026,即ADOPD的面向推理的扩展版本,它将页面分解转化为接地的文档理解。ADOPD 2026在ADOPD 2024数据集继承的页面锚点基础上,补充了人工清洗的标题、语义标签以及与文档区域接地的生成式思维链(CoT)轨迹。我们没有将边界框、掩码和标签视为独立的监督信号,而是将文本块、视觉实体、语义标签、边界框和多边形掩码视为视觉锚点的共享词汇表。该表示支持三种关联能力:其一,区域级语义标注要求模型从页面上下文和局部外观中识别文档元素类型,揭示了标准布局基准常隐藏的长尾语义失败;其二,统一视觉-语言接地生成文本区域和视觉实体,同时附带坐标或多边形轮廓,将检测和分割输出转化为可被下游推理系统复用的结构化锚点;其三,在源自ADOPD 2026的基准DocCount上评估发现,当前最先进的模型仍在密集计数任务上表现不佳,凸显了文档语义理解中基于锚点思考的流程的必要性。通过将页面分解与可验证的视觉锚点推理相连接,ADOPD 2026提供了一个任务框架,推动文档理解从定位转向基于锚点的文档智能。
英文摘要
Existing document understanding benchmarks have largely focused on locating page elements, yet real-world document intelligence requires models to reason jointly about region semantics, spatial relations, and visual structure. We present ADOPD 2026, a reasoning-oriented extension of ADOPD that turns page decomposition into spatially grounded document understanding. ADOPD 2026 enriches page anchors inherited from ADOPD 2024 dataset with human-cleaned captions, semantic tags, and generated chain-of-thought (CoT) traces grounded to document regions. Instead of treating boxes, masks, and tags as independent supervision signals, we cast text blocks, visual entities, semantic labels, bounding boxes, and polygon masks as a shared vocabulary of visual anchors. This representation supports three connected capabilities. First, region-level semantic tagging asks models to identify document element types from both page context and local appearance, revealing long-tail semantic failures that standard layout benchmarks often hide. Second, unified vision-language grounding generates text regions and visual entities together with coordinates or polygonal outlines, transforming detection and segmentation outputs into structured anchors that can be reused by downstream reasoning systems. Third, current state-of-the-art models still struggle with dense counting tasks evaluated on DocCount, a benchmark derived from ADOPD 2026, highlighting the need for the Thinking-with-Anchors pipeline in document semantic understanding. By connecting page decomposition to verifiable visual-anchor reasoning, ADOPD 2026 provides a task framework that moves document understanding beyond localization toward anchor-grounded document intelligence.