发表机构
Shanghai Artificial Intelligence Laboratory; Shanghai Jiao Tong University(上海人工智能实验室; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出用于评估多模态大语言模型科学指令遵循能力的SciMIF基准,揭示模型在科学领域性能差异及规模提升未对应改善约束遵循等问题,填补相关评估空白。
AI 中文摘要
理解科学领域的指令遵循能力,对于有效利用多模态大语言模型(MLLMs)推动科学领域发展至关重要。本研究中,我们提出SciMIF,这是一个用于评估MLLMs遵循复杂科学指令能力的新型基准。具体而言,通过对5个代表性科学领域的22项不同任务进行广泛分析,我们提出了包含10个约束组的综合分类法,该分类法既涵盖了通用功能要求,也包含了学科特定特征。在该分类法的指导下,我们开发了高保真的指令注入流水线,以系统地扩充现有科学数据集。我们对多个最先进的闭源及开源MLLMs开展了全面实验。研究结果显示,不同科学领域间存在显著的性能差异,其中化学领域对当前MLLMs构成更大挑战。此外,我们观察到增大模型规模并未带来约束遵循能力的相应提升,当前模型在细粒度约束及需深度应用学科知识的指令上仍存在严重困难。SciMIF填补了科学领域内多模态指令遵循评估的现有空白,为未来提升MLLMs在严谨科学应用中的性能奠定了关键基础。数据与代码将发布于该httpsURL。
英文摘要
Understanding instruction-following capabilities in scientific domains is essential for effectively leveraging Multimodal Large Language Models (MLLMs) to advance the development of scientific fields. In this work, we introduce SciMIF, a novel benchmark designed to evaluate the capability of MLLMs in following complex scientific instructions. Specifically, based on an extensive analysis of 22 distinct tasks across 5 representative scientific disciplines, we propose a comprehensive taxonomy comprising 10 constraint groups that captures both general functional requirements and discipline-specific characteristics. Guided by this taxonomy, we develop a high-fidelity instruction injection pipeline to systematically augment existing scientific datasets. We conduct comprehensive experiments on multiple state-of-the-art closed-source and open-source MLLMs. Our findings reveal significant performance disparities across different scientific disciplines, with chemistry posing greater challenges for current MLLMs. Furthermore, we observe that increasing the model scale does not yield corresponding improvements in constraint adherence, and current models still struggle severely with fine-grained constraints and instructions requiring the deep application of disciplinary knowledge. SciMIF fills the current void in evaluating multimodal instruction adherence within scientific domains, laying a crucial foundation for future enhancements of MLLMs in rigorous scientific applications. Data and code will be released at https://github.com/shenye7436/SciMIF .
Comments21pages, 9 figures, 16 tables