发表机构
Fudan University; Beijing Institute of Technology(复旦大学; 北京理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有AIVC基准局限于模拟层的问题,提出OmniVCBench基准,通过6,077个图源问答对及AIVC-Judge评估框架,测试多模态模型在解释证据和形成假设方面的能力。
AI 中文摘要
人工智能虚拟细胞(AIVCs)被设想为能够模拟细胞反应、解释潜在机制并支持假设驱动发现的科学智能体。然而,现有的AIVC基准主要在模拟层面运作,这促使我们需要补充评估模型如何解释实验证据并形成生物学假设。我们引入了OmniVCBench,这是一个以图形为中心、来源可追溯的基准,用于评估AIVC的解释组件。它包含从科学文献中的图形和实验背景中提取的6,077个经过筛选的单子图和多子图问答对。在布鲁姆分类法的指导下,我们通过三个科学推理任务实例化了AIVC的预测-解释-发现议程在解释层面的对应任务。我们进一步引入了AIVC-Judge,这是一个任务条件化的多模态大语言模型作为评判者的框架,具有类别特定、参考感知的评分标准,用于评估开放式回答。一种互补的模型派生硬负样本挖掘(MDHNM)策略将模型推理过程中观察到的合理错误转化为多项选择题的干扰项,以降低评估成本。在评估的异构模型池中,多项选择题准确率与AIVC-Judge得分呈正相关,为开放式回答评估提供了补充的性能视角。代码和数据演示可在该https URL获取。
英文摘要
Artificial Intelligence Virtual Cells (AIVCs) are envisioned as scientific agents that simulate cellular responses, explain underlying mechanisms, and support hypothesis-driven discovery. Existing AIVC benchmarks, however, operate primarily at the simulation layer, motivating complementary evaluation of how models interpret experimental evidence and formulate biological hypotheses. We introduce OmniVCBench, a figure-centric, source-traceable benchmark for the interpretation component of an AIVC. It contains 6,077 curated single- and multi-subfigure question--answer pairs derived from figures and experimental contexts in the scientific literature. Guided by Bloom's taxonomy, we instantiate interpretation-layer counterparts of the AIVC Predict--Explain--Discover agenda through three scientific reasoning tasks. We further introduce AIVC-Judge, a task-conditioned MLLM-as-a-judge framework with category-specific, reference-aware rubrics for evaluating open-ended responses. A complementary Model-Derived Hard-Negative Mining (MDHNM) strategy converts plausible errors observed during model inference into MCQ distractors for lower-cost evaluation. Within the evaluated heterogeneous model pool, MCQ accuracy correlates positively with AIVC-Judge scores, providing a complementary view of performance alongside open-response evaluation. Code and data demo are available at https://anonymous.4open.science/r/OmniVCBench.
Comments45 pages, 16 figures;