arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TrialAtlas:用于临床试验设计与优化的多智能体研究组织

TrialAtlas: Multi-Agent Research Organization for Clinical Trial Design and Optimization

Jiacheng Lin, Zifeng Wang, Zheng Chen, Erick Scott, Ziwei Yang, Fanyang Yu, Sheng Zhong, Jimeng Sun

arXiv 2609.21859首次发表:更新:

发表机构

Siebel School of Computing and Data Science, University of Illinois Urbana-Champaign; Keiji AI Inc; Institute of Scientific and Industrial Research, The University of Osaka; University of Pennsylvania; AbbVie Inc; Carle Illinois College of Medicine, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校西贝尔计算与数据科学学院; Keiji AI 公司; 大阪大学产业科学研究所; 宾夕法尼亚大学; 艾伯维公司; 伊利诺伊大学厄巴纳-香槟分校卡尔·伊利诺伊医学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

TrialAtlas通过多智能体协作与记忆增强,优化临床试验设计并预测成败,在FDA数据上显著超越基线。

AI 中文摘要

尽管投入了数十亿美元,仍有近90%进入临床开发的药物最终失败。因此,制药公司依赖临床开发规划(CDP)以及技术与监管成功概率评估来预判开发风险,然而这些决策仍然劳动密集且主观,需要临床科学、统计学、监管事务和竞争情报领域的专家共同获取、综合并推理异构证据。在此,我们引入TrialAtlas,一个用于CDP的、具有记忆增强功能的多智能体研究组织,它通过协调专门负责文献综合、竞争性试验情报、监管先例分析以及试验设计与开发风险综合推理的智能体,来模拟这一协作过程。TrialAtlas还从历史临床试验和监管结果(包括先前的新药申请(NDAs))中学习,以将其决策基于积累的开发经验。为了在真实的监管环境中评估这些能力,我们引入了TrialAtlasBench,该基准由291封FDA完整回应函构建,涵盖三个实际任务:检测试验设计缺陷、推荐可操作的改进设计以及预测技术与监管成功。TrialAtlas在缺陷检测上取得了50.0%的F1分数,比最强基线高出6.1个百分点;在技术与监管成功预测上达到了85.3%的平衡准确率和84.7%的F1分数,在平衡准确率上比最佳基线提高了6.7个百分点,在Cohen's kappa上提高了12.0个百分点。在专家评估中,TrialAtlas生成的关注点中有86.4%被判定为有效,而OpenAI DeepResearch为83.1%,Gemini DeepResearch为59.3%。

英文摘要

Nearly 90% of drugs entering clinical development ultimately fail, despite billions of dollars in investment. Pharmaceutical companies therefore rely on clinical development planning (CDP) and probability of technical and regulatory success assessment to anticipate development risks, yet these decisions remain labor-intensive and subjective, requiring experts across clinical science, statistics, regulatory affairs, and competitive intelligence to jointly acquire, synthesize, and reason over heterogeneous evidence. Here, we introduce TrialAtlas, a memory-augmented multi-agent research organization for CDP that mirrors this collaborative process by coordinating specialized agents for literature synthesis, competitive trial intelligence, regulatory precedent analysis, and integrated reasoning over trial design and development risk. TrialAtlas further learns from historical clinical trials and regulatory outcomes, including prior New Drug Applications (NDAs), to ground its decisions in accumulated development experience. To evaluate these capabilities in an authentic regulatory setting, we introduce TrialAtlasBench, constructed from 291 FDA Complete Response Letters and spanning three practical tasks: detecting trial design deficiencies, recommending actionable design improvements, and predicting technical and regulatory success. TrialAtlas achieves an F1 score of 50.0% for deficiency detection, outperforming the strongest baseline by 6.1 points, and reaches 85.3% balanced accuracy and 84.7% F1 for prediction of technical and regulatory success, improving over the best baselines by 6.7 points in balanced accuracy and 12.0 points in Cohen's kappa. In expert evaluation, 86.4% of TrialAtlas-generated concerns were judged valid, compared with 83.1% for OpenAI DeepResearch and 59.3% for Gemini DeepResearch.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑