arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03606cs.AIcs.LG

临床试验策略学习:面向决策智能体的离线策略训练

Learning Clinical-Trial Strategy: Offline Policy Training for Decision Agents

William Bolton, Philip Torr

首次发表
浏览论文内容

中文总结 AI 辅助

本研究将肿瘤学临床开发建模为离线决策问题,构建时间数据集,对比4种离线目标与4种LLM智能体,发现奖励加权行为克隆表现最优,证明结构化离线学习可用于规划临床实验。

中文摘要 AI 辅助

临床开发是不确定性下的序列决策问题,申办方必须基于异质性证据规划试验组合。本研究将肿瘤学临床开发建模为离线决策问题,智能体需根据决策日可用信息预测肿瘤药物项目未来6个月的试验组合。为此,构建时间数据集,整合31700条异质性公共数据记录(含试验注册、监管审查、申办方申报、使用数据及流行病学数据),形成45个历史项目的881个离线决策回合。对比4种离线目标(行为克隆、奖励加权行为克隆、学习奖励训练、基于价值的隐式Q学习)与4种前沿大语言模型(LLM)智能体,这些智能体在保留药物、申办方、药物类别及时间划分中共享日期门控检索框架。离线训练的模型优于未微调基线,尤其在2025年8月后污染清理保留集中表现突出;奖励加权行为克隆表现最佳,指标指示F1达46.2%、严格F1达14.2%,而对应指标表现最优的工具智能体仅为25.0%和2.1%。这些结果表明,结构化离线学习可教会智能体规划临床实验。

英文摘要

Clinical development is sequential decision-making under uncertainty, where a sponsor must plan a portfolio of experiments from heterogeneous evidence. We study this setting by framing oncology clinical development as an offline decision-making problem in which an agent predicts the next six-month trial portfolio of an oncology drug program from information available at the decision date. To support this, we construct a temporal dataset that combines 31.7k heterogeneous public data records, including trial registries, regulatory reviews, sponsor filings, utilization data, and epidemiology, into 881 offline decision episodes across 45 historical programs. We compare four offline objectives: behavioral cloning, reward-weighted behavioral cloning, learned-reward training, and value-based implicit Q-learning against four frontier LLM agents that share a common date-gated retrieval scaffold across held-out drug, sponsor, drug-class, and temporal splits. Models trained offline outperform the non-fine-tuned baselines, particularly in the post-August 2025 contamination-clean holdout. Reward-weighted behavioral cloning performs the best, obtaining 46.2% indication F1 and 14.2% strict F1 against 25.0% and 2.1%, respectively, for the best-performing tool agent on each metric. These results suggest that structured offline learning can teach agents to plan clinical experiments.

补充信息

↑