发表机构
Shanghai Academy of AI for Science(上海人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对科学多领域推理需求,引入统一科学多模态模型MKB,围绕共享Transformer主干及定制组件构建,涵盖多科学分支。通过两阶段训练,在多方面实验表现出色,证明该范式可行,为跨领域科学多模态探索提供基础。
AI 中文摘要
科学发现正从孤立学科转向多领域推理,人工智能也面临类似转变。现有系统要么专注单一领域,要么主要通过文本分词和基于提示的接口统一科学数据,限制了处理多样科学输入、生成模态原生输出及跨领域联合理解、推理和生成的能力。我们引入MKB,一个围绕共享Transformer主干及模态定制的编码器、适配器和解码器构建的用于理解和生成的统一科学多模态模型。MKB涵盖六个科学分支,支持多种原生输出。训练采用两阶段课程:第一阶段将特定模态组件与冻结主干对齐,第二阶段使用混合科学和通用语料库与语言主干整合。实验表明,MKB在生物和分子基准测试中实现了有竞争力的科学理解,在天气预报、生物生成和医学图像分割方面产生了高保真原生输出,并很大程度保留了Qwen3-VL主干的通用能力。这些结果证明了所提范式的可行性,表明具有模态定制组件的共享主干模型可为未来跨领域科学多模态探索提供有前景的基础。模型和代码可公开获取。
英文摘要
Scientific discovery is increasingly shifting from isolated disciplines to multi-domain reasoning, and AI for science faces a similar transition. Existing systems are either specialised for individual domains or unify scientific data mainly through text tokenisation and prompt-based interfaces, limiting their ability to handle diverse scientific inputs, produce modality-native outputs, and support joint understanding, reasoning, and generation across scientific domains. We introduce MKB, a unified scientific multimodal model for both understanding and generation, built around a shared Transformer backbone and modality-tailored encoders, adapters, and decoders. MKB covers six scientific branches, including DNA, RNA, proteins, small molecules, earth science, and medical images, and supports native outputs such as biological sequences, molecular strings, meteorological fields, and segmentation masks. Training follows a two-stage modality-then-language curriculum: Stage 1 aligns modality-specific components with the frozen backbone, and Stage 2 consolidates them with the language backbone using mixed scientific and general corpora. Experiments show that MKB achieves competitive scientific understanding across biological and molecular benchmarks, produces high-fidelity native outputs for weather forecasting, biological generation, and medical-image segmentation, and largely retains the general capabilities of its Qwen3-VL backbone. These results demonstrate the feasibility of the proposed paradigm, suggesting that shared-backbone models with modality-tailored components can provide a promising foundation for future cross-domain scientific multimodal exploration. The model and code are publicly available at https://github.com/Shanghai-Academy-of-AI-For-Science/MKB and https://huggingface.co/sais-org/MKB.