OphIn-500K:策划网络规模的视觉指令以扩展眼科多模态大语言模型
OphIn-500K: Curating Web-Scale Visual Instructions for Scaling Ophthalmic Multimodal Large Language Models
- Arizona State University(亚利桑那州立大学)
- Clemson University(克莱姆森大学)
- Washington University in St. Louis(圣路易斯华盛顿大学)
- University of Notre Dame(诺特丹大学)
- Florida State University(佛罗里达州立大学)
- Rice University(里德大学)
- NVIDIA(英伟达)
- Mayo Clinic(梅奥诊所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出OphIn-Engine流水线从网络视频中构建高质量眼科指令数据,生成包含50万+指令实例的OphIn-500K数据集,并基于此开发眼科专用多模态大语言模型OphIn-VL,在多项任务上超越现有通用医学和专用模型。
AI中文摘要:
通用医学多模态大语言模型(MLLMs)的进步为构建支持临床诊断的对话助手展现了巨大潜力。然而,它们在高度专业化领域(如眼科)的适应性仍未得到充分探索,主要原因是缺乏大规模、领域特定的指令微调数据。现有的眼科对话数据集通常规模有限,且大多依赖于已建立的公共基准图像,限制了眼科MLLMs的可扩展性及其捕捉真实临床复杂性的能力。为解决这一问题,我们提出了$ extbf{OphIn-Engine}$,一个眼科特定的指令数据策划流水线,从开放获取的眼科网络规模视频中构建高质量指令数据。该流水线整合了多模态转录以提取图像-文本对、视觉线索分离与评分以识别临床相关的视觉描述,以及指令合成与质量控制以生成准确且多样的临床对话。利用该引擎,我们推出了$ extbf{OphIn-500K}$,一个大规模多模态眼科指令微调数据集,包含超过50万个指令实例和来自29,000多个视频片段的151,000多张独特图像,格式包括视觉问答(VQA)、多轮对话交互和思维链(CoT)推理。基于该数据集,我们进一步开发了$ extbf{OphIn-VL}$,一个具有高级视觉理解和对话能力的眼科专用MLLM。综合实验和案例研究表明,与最先进的通用医学和领域专用MLLMs相比,OphIn-VL实现了更优的性能。
英文摘要:
The advancement of general medical Multimodal Large Language Models (MLLMs) has shown great potential for building conversational assistants to support clinical diagnosis. However, their adaptation to highly specialized domains such as ophthalmology remains underexplored, primarily due to the scarcity of large-scale, domain-specific instruction-tuning data. Existing ophthalmic datasets for conversational agents are often limited in scale and largely rely on images from established public benchmarks, limiting the scalability of ophthalmic MLLMs and their ability to capture real-world clinical complexity. To address this gap, we propose $\textbf{OphIn-Engine}$, an ophthalmology-specific instruction data curation pipeline that constructs high-quality instruction data from open-access ophthalmology web-scale videos. The pipeline integrates multimodal transcription for extracting image-transcript pairs, visual cue separation and scoring for identifying clinically relevant visual descriptions, and instruction synthesis with quality control for generating accurate and diverse clinical dialogues. Using this engine, we introduce $\textbf{OphIn-500K}$, a large-scale multimodal ophthalmology instruction-tuning dataset containing over 500,000 instruction instances and more than 151,000 unique images from over 29,000 video clips, formatted as visual question answering (VQA), multi-turn conversational interactions, and chain-of-thought (CoT) reasoning. Built upon this dataset, we further develop $\textbf{OphIn-VL}$, an ophthalmology-specific MLLM with advanced visual understanding and conversational capabilities. Comprehensive experiments and case studies demonstrate that OphIn-VL achieves superior performance compared with state-of-the-art general medical and domain-specific MLLMs.