arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语言模型对抑郁的评分反映的是评分者而非患者

Language-model ratings of depression reflect the rater more than the patient

Baihan Lin

arXiv 2610.08501首次发表:更新:

发表机构

Icahn School of Medicine at Mount Sinai; James J. Peters VA Medical Center; Harvard University(西奈山伊坎医学院; 詹姆斯·J·彼得斯退伍军人事务医疗中心; 哈佛大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过880个语言模型评分者评估抑郁症访谈,发现模型选择主导评分差异,校准虽提升准确率但无法消除个体层面的评分者分歧。

AI 中文摘要

抑郁症没有诊断性的血液检测。语言模型有望提供不知疲倦、一致的评估,但准确的评分者会对个体产生分歧吗?我们预先注册了880个语言模型评分者,将11个开放模型与提示和评分选择交叉组合,并将其应用于189个访谈,以八项患者健康问卷为基准。模型选择解释了30.0%的症状总分方差,稳定的参与者差异解释了10.5%。两个随机抽取的评分者(其受试者工作特征曲线下面积(AUC)≥0.70)平均对40%的参与者的筛查决定存在分歧。平均过度评分决定了被标记的人数,但同等能力的评分者对约五分之一的参与者做出了不同的选择。对86个新访谈的锁定分析重现了主要的预先注册发现。使用40个标记参与者的探索性重新校准将准确率从约60%提高到75%,并将分歧减半,但仍使五分之一的参与者被不同地决定。校准修复了大部分评分者依赖性,但未能确保对个体的共识。

英文摘要

Depression has no diagnostic blood test. Language models promise tireless, consistent assessment, but can accurate raters disagree about individuals? We pre-registered 880 language-model raters, crossing 11 open models with prompting and scoring choices, and applied them to 189 interviews against the eight-item Patient Health Questionnaire. Model choice explained 30.0% of summed-symptom score variance, stable participant differences 10.5%. Two randomly drawn raters with area under the receiver operating characteristic curve (AUC) >= 0.70 disagreed on screening decisions for 40% of participants, on average. Average over-rating governed how many were flagged, yet equal-capacity raters chose differently for about one participant in five. A locked analysis of 86 new interviews reproduced the main pre-registered findings. Exploratory recalibration with 40 labelled participants raised accuracy from about 60% to 75% and halved disagreement, leaving one participant in five decided differently. Calibration repaired much of the rater dependence without securing agreement about individuals.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑