发表机构
TIB - Leibniz Information Centre for Science and Technology; L3S Research Center, Leibniz University(蒂宾根信息中心(莱布尼茨科学技术信息中心); 莱布尼茨大学L3S研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究介绍了ICDAR 2026 ALD/E科学图信息提取竞赛,基于Sci-ImageMiner基准数据集开展多任务评测,发现现有多模态模型在数据提取等任务中存在局限,为领域多模态AI研究提供了平台。
AI 中文摘要
使用多模态人工智能进行科学图理解与推理,需将视觉感知与领域特定推理相结合,以提取研究出版物文本中未呈现的有意义知识。Sci-ImageMiner基准数据集伴随一场社区驱动的竞赛,通过整理涵盖四个端到端互补任务的全面、经专家标注的数据集,超越了以往科学竞赛的水平。该竞赛于2026年1月9日至4月8日期间吸引了68名活跃参与者,共收到1263份公开/提交作品。研究结果显示,最先进的多模态模型在分类和摘要任务中表现良好,但在数据提取和科学推理方面存在困难,尤其是在视觉问答任务中。这些发现揭示了关键局限性,凸显了改进领域感知多模态人工智能系统的挑战与机遇。总体而言,Sci-ImageMiner基准数据集和竞赛建立了一个严谨的平台,以推进科学图理解与推理领域的研究,并展示了最先进方法在这一具有挑战性和复杂性的研究领域的潜力。
英文摘要
Scientific figure comprehension and reasoning using multimodal AI requires integrating visual perception with domain-specific reasoning to extract meaningful knowledge, often not presented in the text of a research publication. The Sci-ImageMiner benchmark dataset, accompanied by a community-driven competition, raises the bar over prior scientific competitions by curating a comprehensive, expert-annotated dataset across four end-to-end complementary tasks. The competition attracted 68 active participants and 1,263 public/private submissions from 9th January 2026 to 8th April 2026. Our results show that state-of-the-art multimodal models perform well on classification and summarization tasks but struggle with data extraction and scientific reasoning, particularly in visual question-answering. These findings reveal key limitations and highlight challenges and opportunities for improving domain-aware multimodal AI systems. Overall, the Sci-ImageMiner benchmark and competition establish a rigorous platform for advancing research in scientific figure comprehension and reasoning and demonstrate the potential of state-of-the-art approaches for a challenging and complex research area.
Comments17 pages, 6 figures, 8 tables, to be published in ICDAR 2026 Conference Proceedings and Volume 16975 of the Lecture Notes in Computer Science series (Springer Nature)