arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GxP-Agent:基于流程有向无环图拓扑结构的可靠临床试验编程大模型智能体

GxP-Agent: Process-DAG Topology for Reliable Clinical Trial Programming with LLM Agents

Jaime Yan

arXiv 2608.16890首次发表:更新:

发表机构

Harrisburg University of Science and Technology(哈里斯堡科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出基于流程DAG拓扑的多智能体系统GxP-Agent,在临床试验编程任务上显著优于现有方法,实现了高结构匹配率,验证了编码领域流程知识为图拓扑的有效性。

AI 中文摘要

临床试验编程是指将研究方案转换为符合CDISC标准、可用于分析的数据集,是监管申报的瓶颈环节,但基于大语言模型(LLM)的代码生成在该任务上表现极差:对5种前沿模型各进行11次单样本尝试,均未生成有效的受试者层面分析数据集。我们提出GxP-Agent,这一多智能体系统将监管流程的顺序编码为有向无环图(DAG),将整体数据集生成分解为15个领域特定节点,由具备pharmaverse技能上下文、验证门控和条件重试机制的工作智能体执行。在基于FDA试点申报CDISCPilot01构建的新执行基准CDISC-Bench(含254名受试者、49个真实ADSL变量)上,配备Claude Sonnet 4.6的GxP-Agent在3次独立运行中均达到100%结构匹配(49个变量全部正确,254条记录全部正确),而最佳检索增强基准的结构匹配率为59.2%,所有单智能体及扁平多智能体方法的结构匹配率均为0%。该DAG拓扑结构还能让性能较弱的模型发挥作用:在相同DAG下,GPT-4.1的平均结构匹配率达59.2%,而在其他所有架构下其得分均为0%。该方法还可推广至ADAE(不良事件数据集,含9节点分支DAG、55个变量、1191条记录),首次尝试即达到100%结构匹配。这些结果表明,将领域流程知识编码为图拓扑结构而非仅依赖LLM推理,是实现符合GxP标准的可靠临床试验编程的关键推动因素。

英文摘要

Clinical trial programming -- transforming study protocols into analysis-ready datasets under CDISC standards -- is a bottleneck in regulatory submissions, yet LLM-based code generation fails catastrophically on this task: across 11 single-shot attempts with five frontier models, none produces a valid subject-level analysis dataset. We introduce GxP-Agent, a multi-agent system that encodes regulatory process ordering as a directed acyclic graph (DAG), decomposing monolithic dataset generation into 15 domain-specific nodes executed by worker agents with pharmaverse skill context, validation gates, and conditional retry. On CDISC-Bench, a new execution-based benchmark built from the FDA pilot submission CDISCPilot01 (254 subjects, 49 ground-truth ADSL variables), GxP-Agent with Claude Sonnet 4.6 achieves 100% structural match (49/49 variables, 254 correct records) across three independent runs, compared to 59.2% for the best retrieval-augmented baseline and 0% for all single-agent and flat multi-agent approaches. The DAG topology also enables weaker models: GPT-4.1 achieves 59.2% mean structural match under the same DAG, where it scores 0% under every other architecture. The approach generalizes to ADAE (adverse events; 9-node branching DAG, 55 variables, 1,191 records), achieving 100% structural match on the first attempt. These results demonstrate that encoding domain process knowledge as graph topology -- rather than relying on LLM reasoning alone -- is a key enabler for reliable, GxP-compliant clinical trial programming.

CommentsPreprint. 9 pages main text, 3 figures, plus references and appendix

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑