AI聊天机器人能否找到专家会选择的研究?模型、用户角色和样本量对医学问题研究检索的影响
Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions
浏览论文内容
中文总结 AI 辅助
该研究评估Claude Sonnet 5、Gemini 3.1 Pro、ChatGPT GPT-5.5三款LLM聊天机器人在模拟三种用户角色下的医学研究检索表现,发现其召回率因模型、角色差异显著,且存在偏向大样本临床试验的偏差。
中文摘要 AI 辅助
大型语言模型(LLM)聊天机器人越来越多地被用于回答临床问题,并附上相关临床研究的引用。现有研究主要关注引用伪造问题,却在评估检索研究的质量及其选择驱动因素方面存在空白。本研究评估了三款通用型LLM聊天机器人:Claude Sonnet 5、Gemini 3.1 Pro和ChatGPT GPT-5.5。我们使用改编自2026年《Cochrane系统评价数据库》第6、7期的20个综述问题的临床问题提示模型,模拟患者、临床医生和证据合成研究者三种用户角色。每个聊天机器人在每种用户角色下各进行4次独立重复查询,共产生720条响应。要求各聊天机器人以原始临床引用为答案提供支撑,我们将其与Cochrane综述的纳入和排除研究集进行基准对比。平均而言,聊天机器人响应检索到39.2%±29.8%的Cochrane纳入研究,同时引用5.0%±9.4%的排除研究。Cochrane纳入研究的召回率因模型和用户角色差异显著:ChatGPT的召回率高于Claude或Gemini(63.1%±29.5% vs. 37.0%±23.8% vs. 17.3%±13.1%;p=2.0×10⁻⁵);研究者角色的召回率高于临床医生或患者角色(42.8%±30.8% vs. 38.6%±28.9% vs. 36.1%±29.3%;p=2.0×10⁻⁵)。在控制发表年份、年引用量和开放获取状态后,样本量是检索的唯一独立显著预测因子(对数样本量每增加1单位,优势比为1.80,95%置信区间1.37-2.36,p=2.34×10⁻⁵)。这些发现表明,尽管LLM聊天机器人可检索到专家评审员识别的部分研究,但其性能因模型和用户角色而异,且存在偏向样本量更大的临床试验的偏差。
英文摘要
Large language model (LLM) chatbots are increasingly used to answer clinical questions with citations to relevant studies, yet the quality of retrieved evidence and factors influencing study selection remain unclear. We evaluated three general-purpose LLM chatbots (Claude Sonnet 5, Gemini 3.1 Pro, and ChatGPT GPT-5.5) using 20 clinical questions adapted from 2026 Cochrane reviews. We simulated patient, clinician, and evidence-synthesis researcher roles and obtained four independent responses for each chatbot-role-question combination, yielding 720 responses (3 chatbots $\times$ 3 user roles $\times$ 4 repetitions $\times$ 20 review questions). Chatbots were asked to support their answers with primary clinical citations, which were benchmarked against the included and excluded study sets of the corresponding Cochrane reviews. On average, a single response retrieved 39.2% $\pm$ 29.8% of the corresponding Cochrane included-study set and 5.0% $\pm$ 9.4% of the excluded-study set. Recall of included studies varied significantly by model and user role. ChatGPT achieved higher recall than Claude or Gemini (63.1% $\pm$ 29.5% vs. 37.0% $\pm$ 23.8% vs. 17.3% $\pm$ 13.1%; blocked permutation test, $p=2.0\times10^{-5}$), and the researcher role yielded higher recall than the clinician or patient roles (42.8% $\pm$ 30.8% vs. 38.6% $\pm$ 28.9% vs. 36.1% $\pm$ 29.3%; $p=2.0\times10^{-5}$). Controlling for publication year, citations per year, and open-access status, sample size was the only significant predictor of retrieval: each doubling of sample size was associated with 50% higher odds of retrieval (odds ratio 1.50, 95% CI 1.24-1.81). These findings show that LLM chatbots can retrieve studies identified by expert reviewers, but retrieval varies substantially across models and user roles and favors larger clinical trials.
发表机构
- National Institute on Drug Abuse(国家药物滥用研究所)
- National Institutes of Health(美国国立卫生研究院)
- National Library of Medicine(国家医学图书馆)
- School of Information Sciences, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校信息科学学院)
机构由 AI 辅助整理,请以论文原文为准。