arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DrugTargetWorld:用于训练和基准测试AI科学家的合成生物库

DrugTargetWorld: A Synthetic Biobank for Training and Benchmarking AI Scientists

Samuel Margolis, Paul Schmiedmayer, Alan Huang, Ethan Chen, Ishan Bhattacharjee, Atman Shah, Ben Viggiano, Fang Cao, Shriya Reddy, Roger Xia, Jack O'Sullivan, Daniel Katz, Matthew Wheeler, Euan Ashley, Bruna Gomes

arXiv 2610.09558首次发表:更新:

发表机构

Stanford University; Brown University(斯坦福大学; 布朗大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DrugTargetWorld通过程序化生成具有隐藏因果结构的模拟生物库,将药物靶点发现转化为可验证的AI训练与评估问题,但现有代理在区分非因果蛋白上仍存在局限。

AI 中文摘要

药物靶点发现需要区分那些因果驱动疾病的分子与仅与之相关的分子。由于真实世界的生物库缺乏已知的因果真值,且参与者级别的数据受到访问控制,因此端到端地训练和评估AI代理执行这一工作流程十分困难。我们引入了DrugTargetWorld,这是一个程序化生成模拟生物库(即“世界”)的框架,这些生物库具有已知但隐藏的因果结构。每个世界包含54,000名参与者的基因型、蛋白质、健康记录、结局以及合成磁共振成像(MRI)数据。代理必须构建疾病表型,识别因果驱动蛋白,推断调节的有益方向,并可选择性地进行虚拟“湿实验室”实验。我们在20个心血管世界和三种实验预算下,对九个代理进行了540个回合的评估。Opus 5和GPT-5.6 Sol取得了最高的平均综合得分,分别为100分中的39.98和35.38,并且两者平均恢复了64%的因果驱动因素。然而,没有代理能够可靠地区分具有误导性的非因果蛋白,且性能仍受限于连接表型构建、因果证据和干预决策所需的综合判断。通过使每个世界的因果结构对评估者已知但对代理隐藏,DrugTargetWorld将端到端的药物靶点发现转化为一个具有可验证奖励的可扩展训练和评估问题。

英文摘要

Drug target discovery requires distinguishing molecules that causally drive disease from those that are merely associated with it. Training and evaluating AI agents to perform this workflow end-to-end is difficult because real world biobanks lack known causal ground truth and participant-level data is access controlled. We introduce DrugTargetWorld, a framework that procedurally generates simulated biobanks, or "worlds," with known but concealed causal structure. Each world contains genotypes, proteins, health records, outcomes, and synthetic magnetic resonance imaging (MRI) for 54,000 participants. Agents must construct a disease phenotype, identify causal driver proteins, infer the beneficial direction of modulation, and optionally conduct virtual 'wet lab' experiments. We evaluated nine agents in 540 episodes across 20 cardiovascular worlds and three experimental budgets. Opus 5 and GPT-5.6 Sol achieved the highest mean composite scores, 39.98 and 35.38 of 100, respectively, and both recovered 64% of causal drivers on average. However, no agent reliably distinguished misleading non-causal proteins, and performance remained limited by the integrative judgments required to connect phenotype construction, causal evidence, and intervention decisions. By making each world's causal structure known to the evaluator but hidden from the agent, DrugTargetWorld turns end-to-end drug target discovery into a scalable training and evaluation problem with verifiable reward.

Comments34 pages main text, 92 pages supplementary material; 6 main figures. Project: https://DrugTargetWorld.vercel.app. Code: https://github.com/sammargolis/DrugTargetWorld. Data: https://huggingface.co/datasets/sammargolis/DrugTargetWorld-assets

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑