发表机构
University of Texas at Austin(德克萨斯大学奥斯汀分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对大语言模型生成的临床试验摘要的忠实性问题,引入基准评估框架,用特定数据库试验、提示模板和注释模式评估,对比多个模型的基线测量,开发知识图谱增强检索系统,结果显示该系统改进显著,且不同模型改进途径有别。
AI 中文摘要
大语言模型越来越多地用于为医疗保健提供者、患者和付款人总结临床试验结果,但在这种高风险背景下,它们产生幻觉的倾向带来了重大风险。本研究引入了一个基准评估框架,用于衡量大语言模型生成的临床试验摘要在三个利益相关方受众中的忠实性。该框架由从Aggregate Analysis of this http URL数据库中抽取的200个分层试验组成,使用针对特定受众的提示模板和六维忠实性注释模式进行评估。为GPT-4o、Claude Sonnet 4.6和Gemini 2.5 Flash建立了基线测量,在1800个生成的摘要上使用交叉编码器自然语言推理(NLI)模型进行评分。发现无根据的声明是所有三个模型的主要失败模式,平均注释分数为1.55(满分3分)。开发了一个知识图谱增强检索系统并与基线进行评估,在基于NLI的忠实性分数上产生了统计学上的显著改进(蕴含+0.0125,忠实性+0.0130,p < 0.0001)。改进途径因模型而异,GPT-4o主要通过减少矛盾来改进,而Claude Sonnet 4.6和Gemini 2.5 Flash则通过增加蕴含来改进。
英文摘要
Large language models are increasingly used to summarize clinical trial results for healthcare providers, patients, and payers, but their tendency to hallucinate poses significant risks in this high-stakes context. This study introduces a benchmark evaluation framework for measuring the faithfulness of LLM-generated clinical trial summaries across three stakeholder audiences. The framework consists of 200 stratified trials drawn from the Aggregate Analysis of ClinicalTrials.gov database, evaluated using audience-specific prompt templates and a six-dimension faithfulness annotation schema. Baseline measurements were established for GPT-4o, Claude Sonnet 4.6, and Gemini 2.5 Flash across 1,800 generated summaries scored using a cross-encoder natural language inference (NLI) model. Unsupported Claims was identified as the dominant failure mode across all three models, with a mean annotation score of 1.55 out of three. A knowledge-graph-augmented retrieval system was developed and evaluated against the baseline, producing statistically significant improvements in NLI-based faithfulness scores (entailment +0.0125, faithfulness +0.0130, p < 0.0001). Improvement pathways were model-dependent, with GPT-4o improving primarily through contradiction reduction while Claude Sonnet 4.6 and Gemini 2.5 Flash improved through increased entailment.
Comments8 pages, 8 figures