PertReason:用于细胞状态条件下扰动效应机制推理的知识基础基准和框架
PertReason: A Knowledge-Grounded Benchmark and Framework for Cell-State-Conditioned Mechanistic Reasoning of Perturbation Effects
浏览论文内容
中文总结 AI 辅助
研究针对科学领域机器学习评估难题,引入PertReason基准和框架套件,通过PertReasonQA测试模型,结合多数据与知识图谱,发现现有模型预测与推理差距,提出PertReasonLM,为科学系统中忠实推理失败提供诊断框架。
中文摘要 AI 辅助
在科学领域评估机器学习需要在现实分布变化下区分正确预测和正确原因。我们引入了PertReason,这是一个用于细胞状态条件下扰动效应推理的知识基础基准和框架套件。核心的PertReasonQA基准测试模型能否生成机制上可靠的解释,同时对新细胞和未见扰动等复杂变化保持稳健。它结合了多细胞环境下的单细胞遗传和化学扰动数据与知识图谱,动态根据细胞特定基础状态调整通路以避免泛化记忆。对现有模型的评估揭示了预测准确性和机制推理之间的系统差距。作为基准的参考探针,我们提出了PertReasonLM,一个经训练使结果预测与特定情境机制推理对齐的大语言模型。我们共同提供了一个诊断框架,用于揭示和减轻数据丰富的科学系统中忠实推理的失败。
英文摘要
Evaluating machine learning in scientific domains requires separating correct predictions from correct reasons under realistic distribution shifts. We introduce PertReason, a knowledge-grounded benchmark and framework suite for cell-state--conditioned reasoning about perturbation effects. At its core, PertReasonQA is a benchmark that tests whether models can generate mechanistically faithful explanations while remaining robust to complex shifts, such as new cells and unseen perturbations. PertReasonQA combines single-cell genetic and chemical perturbation data across multiple cellular contexts with knowledge graphs, and dynamically conditions pathways on cell-specific basal states to avoid generic memorization. Evaluations on state-of-the-art models reveal systematic gaps between predictive accuracy and mechanistic reasoning. Specifically, these models exhibit failure modes largely invisible to standard benchmarks, such as deriving correct answers through flawed logic, ignoring cellular context, and generating directionally inconsistent mechanisms. As a reference probe of the benchmark, we present PertReasonLM, a large language model trained to align outcome predictions with context-specific mechanistic reasoning. Our model targets the identified failure modes by grounding rationales in context-specific pathways and tightening agreement between outcomes and mechanisms. Together, we provide a diagnostic framework for exposing and mitigating failures in faithful reasoning in data-rich scientific systems.
发表机构
- Texas A&M University(德克萨斯农工大学)
机构由 AI 辅助整理,请以论文原文为准。