发表机构
Icahn School of Medicine at Mount Sinai; James J. Peters VA Medical Center; Harvard University(西奈山伊坎医学院; 詹姆斯·J·彼得斯退伍军人事务医疗中心; 哈佛大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过880个语言模型评分者评估抑郁症访谈,发现模型选择主导评分差异,校准虽提升准确率但无法消除个体层面的评分者分歧。
AI 中文摘要
抑郁症没有诊断性的血液检测。语言模型有望提供不知疲倦、一致的评估,但准确的评分者会对个体产生分歧吗?我们预先注册了880个语言模型评分者,将11个开放模型与提示和评分选择交叉组合,并将其应用于189个访谈,以八项患者健康问卷为基准。模型选择解释了30.0%的症状总分方差,稳定的参与者差异解释了10.5%。两个随机抽取的评分者(其受试者工作特征曲线下面积(AUC)≥0.70)平均对40%的参与者的筛查决定存在分歧。平均过度评分决定了被标记的人数,但同等能力的评分者对约五分之一的参与者做出了不同的选择。对86个新访谈的锁定分析重现了主要的预先注册发现。使用40个标记参与者的探索性重新校准将准确率从约60%提高到75%,并将分歧减半,但仍使五分之一的参与者被不同地决定。校准修复了大部分评分者依赖性,但未能确保对个体的共识。
英文摘要
Depression has no diagnostic blood test. Language models promise tireless, consistent assessment, but can accurate raters disagree about individuals? We pre-registered 880 language-model raters, crossing 11 open models with prompting and scoring choices, and applied them to 189 interviews against the eight-item Patient Health Questionnaire. Model choice explained 30.0% of summed-symptom score variance, stable participant differences 10.5%. Two randomly drawn raters with area under the receiver operating characteristic curve (AUC) >= 0.70 disagreed on screening decisions for 40% of participants, on average. Average over-rating governed how many were flagged, yet equal-capacity raters chose differently for about one participant in five. A locked analysis of 86 new interviews reproduced the main pre-registered findings. Exploratory recalibration with 40 labelled participants raised accuracy from about 60% to 75% and halved disagreement, leaving one participant in five decided differently. Calibration repaired much of the rater dependence without securing agreement about individuals.