面向长文档的查询驱动多模态信息抽取
Query-Driven Multimodal Information Extraction from Long Documents
- School of Artificial Intelligence, Jilin University(吉林大学人工智能学院)
- Graduate School of Information Science and Technology, The University of Tokyo(东京大学信息科学与技术研究生院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对现有多模态长文档信息抽取范式的不足,本文提出查询驱动的图像-文本联合抽取任务,构建基准ITJoint并设计多智能体框架Q2IT,实验显示Q2IT性能显著优于独立VLMs但仍有提升空间。
AI中文摘要:
在特定领域的多模态长文档中,图像与文本共同传递了仅靠纯文本无法完全捕捉的复杂知识。然而,现有的DocVQA等范式主要关注生成文本答案或定位证据区域,而非输出查询特定的文本属性值及对应图像。为解决这一缺口,我们提出面向长文档的查询驱动图像-文本联合抽取任务,要求模型输出查询请求的文本属性值和对应的图像边界框。基于用户意图与文档内容相关的挑战,我们设计了在查询和实例层面运行的两级分类体系。此外,我们构建了首个高质量人工标注的该新任务基准ITJoint,包含2455页特定领域文档、大量非装饰性图像、316个查询及910个答案实例。最后,我们评估了不同厂商的代表性独立视觉-语言模型(VLMs),并进一步设计了Q2IT——一个由三个逐步协作智能体组成的多智能体框架,分别用于证据收集、页面选择和目标图像定位。采用同时评估文本抽取和图像定位的联合评估方法,实验表明独立VLMs在该任务上表现不佳,而Q2IT在ITJoint上的性能显著提升,但距离完美结果仍存在较大差距。
英文摘要:
In domain-specific multimodal long documents, images and text jointly convey complex knowledge that cannot be fully captured by plain text alone. However, existing paradigms like DocVQA primarily focus on generating textual answers or localizing evidence regions, rather than outputting query-specific textual attribute values and corresponding images. To address this gap, we propose query-driven image-text joint extraction from long documents, requiring models to output query-requested textual attribute values and corresponding image bounding boxes. Based on challenges related to both user intent and document content, we designed a two-level taxonomy that operates at the query and instance levels. Further, we construct ITJoint, the first high-quality, manually annotated benchmark for this new task, comprising 2,455 pages of domain-specific documents with numerous non-decorative images, 316 queries, and 910 answer instances. Finally, we evaluate representative standalone Vision-Language Models from different providers and further design Q2IT, a multi-agent collaborative framework consisting of three progressively collaborating agents for evidence collection, page selection, and target-image localization. Using a joint evaluation approach that assesses both text extraction and image localization, our experiments show that standalone VLMs struggle with this task, while Q2IT significantly improves performance on ITJoint, although a substantial gap remains toward perfect results.