arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

仅靠提示是不够的:儿科诊疗场景中利用大语言模型(LLM)衡量共同决策(SDM)的有监督基线与泄漏控制

Prompting is not enough: supervised baselines and leakage control for measuring shared decision-making with LLMs in pediatric encounters

Bernardo Modenesi, Jody Lin, Kimberly Kaphingst, Angela Zhu, Maya Wheeler, Peilu Zhang, Angela Fagerlin

arXiv 2608.14792首次发表:更新:

发表机构

University of Utah; University of Utah Spencer Fox Eccles School of Medicine; Kahlert School of Computing, University of Utah; Huntsman Cancer Institute, University of Utah(犹他大学; 犹他大学斯宾塞·福克斯·埃克尔斯医学院; 犹他大学卡勒特计算机学院; 犹他大学亨茨曼癌症研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对儿科外科决策诊疗场景,对比零样本LLM、有监督分类器及二者的堆叠模型衡量SDM行为的效果,发现零样本提示不足,还识别出数据泄漏路径,为相关评估提供了基线与泄漏控制依据。

AI 中文摘要

目的:确定大语言模型(LLM)的零样本提示是否足以检测真实临床诊疗中的共同决策(SDM)行为,以及在按患者分组的嵌套评估下,有监督学习是否能带来额外价值。方法:我们分析了21份儿童多长期病症家庭与外科医生之间的门诊外科决策诊疗录音(涉及19名独特患者,共7566个话语片段,时长约6.1小时)。经训练的编码员对12种SDM行为的片段进行标注,人类标注者间的宏观Cohen's kappa值为0.695。我们对比了零样本本地LLM(Qwen 2.5 32B)、基于冻结句子嵌入的有监督分类器,以及二者的逻辑回归堆叠模型,评估设置为按患者分组的外层折叠,搭配内层交叉拟合阈值与患者重抽样置信区间。结果:零样本LLM的宏观kappa值为0.139(95%置信区间0.111-0.164);有监督分类器的kappa值为0.227(0.186-0.262),配对提升值为0.088(0.051-0.119);二者的逻辑回归堆叠模型的kappa值为0.242(0.198-0.284)。我们识别出多种语料库特定的泄漏路径,包括将同胞录音单独分组,以及让外层保留患者的标签进入下游模型拟合时使用的少样本示例中。结论:仅靠零样本提示不足以可靠衡量SDM行为,其效果不如小型有监督模型;仅按患者分组无法防止泄漏,当带标签的提示示例在外层评估循环之外预先计算时会出现此类问题。报告的性能对数据拆分单元和带标签示例进入流程的位置敏感,这些发现需外部验证才能推广至本研究之外的人群、模型、提示及编码手册。

英文摘要

Objectives: To determine whether zero-shot prompting of a large language model (LLM) is sufficient to detect shared decision-making (SDM) behaviors in real clinical encounters, and whether supervised learning adds value under patient-grouped, nested evaluation. Methods: We analyzed 21 audio-recorded outpatient surgical decision encounters (19 unique patients; 7,566 utterance segments; ~6.1 hours) between families of children with multiple long-term conditions and their surgical providers. Trained coders labeled segments for 12 SDM behaviors (human-human macro Cohen's kappa = 0.695). We compared a zero-shot local LLM (Qwen 2.5 32B), a supervised classifier over frozen sentence embeddings, and their logistic stack, under patient-grouped outer folds with inner cross-fitted thresholds and patient-resampled confidence intervals. Results: The zero-shot LLM reached macro kappa = 0.139 (95% CI 0.111-0.164). The supervised classifier reached kappa = 0.227 (0.186-0.262), a paired improvement of 0.088 (0.051-0.119). A logistic stack of the two reached kappa = 0.242 (0.198-0.284). We identified multiple corpus-specific leakage paths, including grouping sibling recordings separately and allowing labels from an outer held-out patient to enter few-shot exemplars used while fitting downstream models. Conclusion: Zero-shot prompting alone is not sufficient to measure SDM behavior as reliably as a small supervised model, and patient-level grouping alone does not prevent leakage when labeled prompt exemplars are precomputed outside the outer evaluation loop. Reported performance is sensitive to the unit of data splitting and to where labeled exemplars enter the pipeline. External validation is needed before these findings generalize beyond this population, model, prompt, and codebook.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑