arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25771cs.HCcs.LG

采用困境训练的大语言模型少样本提示在预测患者偏好方面优于人类替代者

Large Language Model Few-Shot Prompting with Dilemma Training Outperforms Human Surrogates in Predicting Patient Preferences

  • National University of Singapore(新加坡国立大学)
  • University of Copenhagen(哥本哈根大学)
  • Imperial College London(伦敦帝国学院)

机构由 AI 辅助整理,请以论文原文为准。

Natasha Ureyang, Sebastian Porsdam Mann, Yuxin Liu, Zuriel Hassirim, Melanie Almonte, Wenhao Chen, Joyce Ng, Thant Nay Lin, Aung Thiha, Gerald CH Koh, Brian Dav… 展开作者

Natasha Ureyang, Sebastian Porsdam Mann, Yuxin Liu, Zuriel Hassirim, Melanie Almonte, Wenhao Chen, Joyce Ng, Thant Nay Lin, Aung Thiha, Gerald CH Koh, Brian David Earp, Pin Sym Foong

AI总结:

该研究提出采用困境训练的P4-DT智能体,通过12组患者-替代者配对实验,其预测患者治疗选择的准确率达81.7%,优于人类替代者及相关方法。

AI中文摘要:

在重症疾病中,人类替代者往往难以准确预测患者偏好(准确率为68%),从而引发决策冲突。个性化患者偏好预测器(P4)智能体提供了潜在解决方案,但此前的原型将价值观视为静态评分,忽略了医疗决策中依赖情境的特性。基于“护理逻辑”,我们提出P4-DT(困境训练),这是一种P4智能体,通过让用户参与不同的医疗困境来构建患者决策策略,通过双向训练引出个体偏好推理。在12组患者-替代者配对的研究中,P4-DT预测患者治疗选择的准确率达81.7%,显著高于随机水平(优势比OR=5.61 [2.03, 15.51],p<0.001),且优于未接受辅助的替代者(准确率55.0%;OR=3.67 [1.59, 8.47],p=0.002)以及接受P4-DT辅助的替代者(准确率61.7%)。对比提示分析显示,与仅使用初始价值观评分相比,纳入情境场景决策和开放式文本可使准确率提高15.0个百分点。我们探讨了进一步测试和设计情境感知AI智能体的意义,这类智能体体现更丰富的人类经验,以在复杂决策中提供协作支持。

英文摘要:

In serious illness, human surrogates often struggle to accurately predict patient preferences (68% accuracy), causing decision conflict. Personalized Patient Preference Predictor (P4) agents offer a potential solution, but prior prototypes treat values as static ratings, ignoring the contextual, situation-dependent nature of medical choices. Grounded in the 'logic of care', we present P4-DT (Dilemma Training), a P4 agent that constructs a patient decision policy by engaging users with varied medical dilemmas, eliciting individual preference reasoning through bi-directional training. In a study with 12 patient-surrogate dyads, P4-DT predicted patient treatment choices with 81.7% accuracy, significantly exceeding chance (OR = 5.61 [2.03, 15.51], p < .001) and outperforming both unassisted surrogates (55.0%; OR = 3.67 [1.59, 8.47], p = .002) and surrogates assisted by P4-DT (61.7%). Comparative prompt analyses showed that incorporating contextual scenario decisions and open-ended text improved accuracy by 15.0 percentage points over initial values ratings alone. We discuss implications for further testing and designing of context-aware AI agents that embody richer human experience to partner in complex decision-making.

↑