VIALS:生命科学领域人工产物视觉解释基准
VIALS: A Benchmark for Visual Interpretation of Artifacts in the Life Sciences
- Handshake AI
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
研究针对生命科学人工产物视觉解释的需求,构建含161项任务的VIALS基准,发现前沿视觉语言模型在该任务上存在领域知识与推理局限,凸显AI需具备此类能力以服务生命科学工作流。
中文摘要 AI 辅助
在专业的生命科学工作流程中,科学家通常会解释凝胶印迹、显微镜图像、质粒图谱、流式细胞术图、分子结构等视觉人工产物,以指导研究决策。我们推出VIALS,这是一个包含161项此类解释任务的视觉问答基准,涵盖了生物技术行业实验工作流程中所检查的人工产物类型,而非来自出版物和教科书的精美图表。尽管前沿的视觉语言模型现在能够流畅地描述自然图像,但我们发现它们无法准确解释这些科学图像,这反映出它们在领域知识和特定领域视觉推理能力方面存在局限性。相比之下,具备相关领域专业知识的科学家认为这些视觉解释任务很简单。无法类似地解释此类图像的人工智能在专业生命科学工作流程中的实用性将十分有限,而这类人工产物是科学家进行推理、交流和决策的核心。
英文摘要
In professional life sciences workflows, scientists routinely interpret visual artifacts (gel blots, microscopy images, plasmid maps, flow cytometry plots, molecular structures, ...) to inform research decisions. We introduce VIALS, a visual question-answering benchmark with 161 such interpretation tasks, spanning the types of artifacts examined throughout experimental workflows in the biotech industry (rather than polished figures from publications and textbooks). While frontier vision-language models can now fluently describe natural images, we find that they are unable to accurately interpret these scientific images, reflecting limitations in domain knowledge and domain-specific visual reasoning capabilities. In contrast, scientists with relevant domain expertise find these visual interpretation tasks straightforward. AI that cannot similarly interpret such images will have limited utility in professional life sciences workflows, where such artifacts are central to how scientists reason, communicate, and make decisions.