arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通用科学人工智能的实现路径:科学图像的多模态理解

A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images

Jennifer D'Souza, Fahad Ahmed, Cecilia Andrea Bustamante Andrade, Lina Frolova, Poorani Gnanasambandan, Dilshad Hussain, Muhammad Uzair Khan, Nkembeng Kevin Nkengfoa, Paul Praveen J., Fabio Priante, Sjoerd Franciscus van der Werf, Thomas Frederik Jan van Roeden

arXiv 2608.14075首次发表:更新:

发表机构

TIB Leibniz Information Centre for Science and Technology; Eindhoven University of Technology; Freie Universität Berlin; International Center for Chemical and Biological Sciences, University of Karachi; National University of Sciences and Technology; University of Warwick; PSG College of Technology; Aalto University(莱布尼茨科学与技术信息中心(TIB); 埃因霍温理工大学; 柏林自由大学; 卡拉奇大学国际化学与生物科学中心; 国家科技大学; 华威大学; PSG理工学院; 阿尔托大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出以科学图像多模态理解构建通用科学AI的路径,依托ALD/E-ImageMiner基准与2026年ICDAR竞赛,明确相关任务能力检验维度及长期研究方向,推动可机器执行的科学视觉知识与可验证多模态科学AI发展。

AI 中文摘要

科学图表和表格蕴含着关键的实验证据,但数字图书馆和多模态人工智能系统仍难以对其进行检索与解读。ALD/E-ImageMiner基准与2026年ICDAR原子层沉积/蚀刻科学图表信息抽取竞赛提供了来自205篇出版物的1951幅图表,经专家标注用于分类、数据表抽取、摘要生成及视觉问答。在配套论文集中,我们阐述了该基准如何指导未来科学图像相关挑战的前瞻性视角,探究其任务如何检验从视觉与定量阅读到领域关联推理及证据论证的能力,以及基于Bloom的问题设计如何助力更深入的科学理解。我们提出“从图像中获取科学概念理解”作为长期基准目标,未来方向涵盖更广泛的领域与图表类型、上下文及跨文档综合、假设评估、溯源、不确定性、反事实依据及开放式多模态研究。该视角将2026年ICDAR竞赛与可机器执行的科学视觉知识及可验证的多模态科学人工智能的更广泛议程相连接。

英文摘要

Scientific figures and tables encode essential experimental evidence, yet remain difficult for digital libraries and multimodal AI systems to retrieve and interpret. The ALD/E-ImageMiner benchmark and ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching Scientific Figures provide 1,951 figures from 205 publications, expert-annotated for classification, data table extraction, summarization, and visual question answering. In these companion proceedings, we present a forward-looking perspective on how the benchmark can guide future scientific-image challenges. We examine how its tasks probe capabilities from visual and quantitative reading to domain-grounded reasoning and evidential justification, and how Bloom-informed question design can support deeper scientific understanding. We propose "scientific conceptual understanding from images" as a long-term benchmark objective, with future directions including broader domains and figure types, contextual and cross-document synthesis, hypothesis evaluation, provenance, uncertainty, counterfactual grounding, and open-ended multimodal research. This perspective connects the ICDAR 2026 challenge to a broader agenda for machine-actionable scientific visual knowledge and verifiable multimodal scientific AI.

Comments15 pages, 1 figure

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑