AI 中文总结
本研究对比7种LLMs与医生的预防性护理优先级排序,发现LLMs与医生高度一致但对生活方式干预重视不足,部分模型表现更优,强化效果需价值对齐训练等。
AI 中文摘要
预防性护理服务(PCS)可延长寿命,但医生往往对生活方式调整等高效干预措施重视不足(Zhang等人,《JAMA网络开放》2020年)。本研究评估大语言模型(LLMs)在时间约束下是否能复制并强化医生对PCS的优先级排序。采用Zhang等人验证过的调查,对两名患者分别进行长、短问诊评估,将7种LLMs与历史医生进行对比;生成137个匹配队列人口统计学特征的模拟医生角色,每种模型测试3种提示词。主要结局为与医生排序的一致性,用斯皮尔曼相关系数衡量,以及共识分层一致性(CSA),即LLMs被评为4分及以上的选择中,与医生共识在各同意层级匹配的比例。次要结局包括每项优先选择获得的生命年数(LYGPC)、一致性和选择性。通过让模型在3种信息提示下修改医生排序来评估强化效果,用LYGPC的变化量量化影响。LLMs与医生高度相似(平均斯皮尔曼相关系数为0.83,标准差0.11),在极端同意范围的CSA较高(94%,210次中有197次),但在中等范围CSA较低(21%,140次中有30次),此时它们对生活方式服务的重视不足(8.8% vs 38%被评为4分及以上;P<0.001)。部分模型在LYGPC、一致性和选择性上优于医生;时间约束对医生和LLMs的影响相似,均提高LYGPC和选择性但降低一致性。强化效果因模型而异。当前LLMs复制了医生的时间敏感性和基线优先级排序,同时加剧了对生活方式干预的重视不足;部分模型改善了优先级排序性能,但要实现一致的强化需开展价值对齐训练、明确时间约束表征及前瞻性真实世界验证。
英文摘要
Preventive care services (PCS) extend life, yet physicians often underprioritize highly effective interventions such as lifestyle modifications (Zhang et al., JAMA Network Open 2020). We evaluated whether large language models (LLMs) replicate and augment physician prioritization of PCS under time constraints. Using Zhang et al.'s validated survey with two patients assessed during long and short visits, we compared seven LLMs with historical physicians. We generated 137 simulated physician personas matching cohort demographics and tested three prompts per model. Primary outcomes were concordance with physician rankings, measured by Spearman correlation, and Consensus-Stratified Agreement (CSA), the proportion of LLM selections rated 4 or higher that matched physician consensus across agreement strata. Secondary outcomes included life-years gained per prioritized choice (LYGPC), consistency, and selectiveness. Augmentation was assessed by having models revise physician rankings under three informative prompts, with delta LYGPC quantifying impact. LLMs closely mirrored physicians (mean Spearman = 0.83, SD = 0.11), with high CSA at extreme agreement ranges (94%, 197/210) but low CSA in moderate ranges (21%, 30/140), where they underprioritized lifestyle services (8.8% vs. 38% rated 4 or higher; P < .001). Several models exceeded physicians in LYGPC and consistency while being more selective. Time constraints affected physicians and LLMs similarly, increasing LYGPC and selectiveness but reducing consistency. Augmentation effects varied by model. Current LLMs reproduced physicians' time-sensitivity and base-rate prioritization while exacerbating underprioritized lifestyle interventions. Some models improved prioritization performance, but consistent augmentation will require value-aligned training, explicit time-constraint representation, and prospective real-world validation.