ICDAR2026多领域文档多模态推理竞赛
ICDAR2026 Competition on Multimodal Reasoning over Documents in Multiple Domains
浏览论文内容
中文总结 AI 辅助
本竞赛通过多领域文档视觉问答任务,评估了多种方法,发现最强系统采用结构化证据提取、检索、验证和多组件编排,而非单次提示。
中文摘要 AI 辅助
本报告介绍了ICDAR2026多领域文档多模态推理竞赛的结果。该竞赛旨在通过视觉问答(VQA)任务推动文档理解研究的发展。在以往DocVQA基准的基础上,本次竞赛引入了涵盖八个领域的多样化文档集合上的挑战性推理问题,这些领域包括商业报告、科学论文、幻灯片、海报、地图、漫画、信息图表和工程图纸。竞赛最终收到来自8个团队的20份有效提交,涉及零样本视觉语言模型(VLM)、OCR和解析器增强流水线、智能体检索系统、多智能体集成以及微调多模态模型。结果表明,最强系统超越了单次提示方法,转而依赖跨多个组件的结构化证据提取、检索、验证和编排。
英文摘要
In this report we present results of the ICDAR2026 Competition on Multimodal Reasoning over Documents in Multiple Domains. This competition aimed to advance research in document understanding through the task of Visual Question Answering (VQA). Building upon previous DocVQA benchmarks, this competition introduces challenging reasoning questions over a diverse collection of documents spanning eight domains, including business reports, scientific papers, slides, posters, maps, comics, infographics, and engineering drawings. The competition concluded with 20 valid submissions from 8 teams spanning zero-shot VLMs, OCR and parser-augmented pipelines, agentic retrieval systems, multi-agent ensembles, and fine-tuned multimodal models. The results show that the strongest systems move beyond single-pass prompting and instead rely on structured evidence extraction, retrieval, verification, and orchestration across multiple components.
发表机构
- Computer Vision Center(计算机视觉中心)
- Universitat Autònoma de Barcelona(巴塞罗那自治大学)
机构由 AI 辅助整理,请以论文原文为准。