arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

人工智能原生生物技术公司需要部门吗?人工智能驱动药物研发的公司世界模型基准测试

Do AI-Native Biotechs Need Departments? Benchmarking Company World Models for AI-Driven Drug Development

Yinan Wang

arXiv 2607.18696首次发表:更新:

发表机构

Noah AI Research(诺亚人工智能研究)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究人工智能原生生物技术公司组织模式,提出公司世界模型,用含45个案例的虚拟实验室基准测试,比较多种架构,发现价值转换架构表现优,核心操作原语应是资产到价值状态而非静态组织结构图,不过研究仅限虚拟实验室。

AI 中文摘要

人工智能原生生物技术公司常通过将人类生物技术组织结构图复制到智能体角色来设计。我们主张一种不同的抽象:公司世界模型,定义为具有转换模型、明确价值函数、规划以及跨科学、监管、业务发展、商业、财务和执行约束进行更新的持久资产到价值状态表示。我们引入一个虚拟实验室基准来测试人工智能智能体组织是应模仿部门还是围绕这样的世界模型运作。该基准包含45个回顾性公共信息决策案例,有严格时间限制、隐藏结果、通用模式、自动评分和盲态成对评判。我们比较了人类组织模仿、更强的人类组织模仿增强、人工智能原生资产中心和人工智能原生价值转换架构。价值转换架构是公司世界模型的提示级近似:通过交易、批准、收入和投资仲裁循环更新的实时资产价值记录。在由外部业务发展、监管批准和推出以及收入规则定义的成功函数下,它获得了最高的自动价值转换分数,并且在特定价值的盲态评判中比原始基线更受青睐。压力测试缩小了论断范围:更强的人类基线仍具竞争力,中立评判未显示出强大的价值转换优势。仅使用Codex的机制性消融表明,收入室、交易室和批准室在目标目标下承担着有用的工作。核心发现是目标敏感的:部门可能仍是有用的治理观点,但人工智能原生的核心操作原语应是共享的、预测性的资产到价值状态,而非静态的人类组织结构图。本研究仅为虚拟实验室研究,未确立现实世界中的药物成功、临床益处或收入预测准确性。

英文摘要

AI-native biotechnology companies are often designed by copying human biotech org charts into agent roles. We argue for a different abstraction: a Company World Model, defined as a persistent asset-to-value state representation with transition models, explicit value functions, planning, and updating across scientific, regulatory, BD, commercial, financial, and execution constraints. We introduce a dry-lab benchmark for testing whether AI-agent organizations should mimic departments or operate around such a world model. The benchmark contains 45 retrospective public-information decision cases with strict time cutoffs, hidden outcomes, common schemas, automatic scoring, and blinded pairwise judging. We compare human-org-mimic, stronger human-org-mimic-plus, AI-native asset-centric, and AI-native value-conversion architectures. The value-conversion architecture is a prompt-level approximation of a Company World Model: a Live Asset Value Record updated by Deal, Approval, Revenue, and Investment Arbiter loops. Under a success function defined by external BD, regulatory approval and launch, and revenue discipline, it achieved the highest automatic value-conversion score and was strongly preferred over the original baselines by value-specific blinded judges. Stress tests narrowed the claim: a stronger human baseline remained competitive, and a neutral judge did not show robust value-conversion dominance. Codex-only mechanistic ablations suggest that Revenue Room, Deal Room, and Approval Room carry useful work under the target objective. The central finding is objective-sensitive: departments may remain useful governance views, but the core AI-native operating primitive should be a shared, predictive asset-to-value state rather than a static human org chart. The study is dry-lab only and does not establish real-world drug success, clinical benefit, or revenue prediction accuracy.

CommentsWorking paper. Includes public no-label benchmark cases and dry-lab evaluation artifacts. No wet-lab, patient-level, clinical, regulatory, or investment validation is claimed

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑