arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39049cs.CLcs.AI

结构 vs. 思维链:评估大语言模型在抑郁严重程度中的标准提取

Structure vs. Chain-of-Thought: Evaluating LLM Criteria Extraction for Depression Severity

  • Independent Researcher(独立研究者)

机构由 AI 辅助整理,请以论文原文为准。

Xinkai Chen

AI总结:

本研究比较了LLM直接评分与基于临床标准提取评分在抑郁严重程度评估中的表现,发现标准提取在多数情况下无显著优势,且可能漏检严重病例,但可提高可审计性。

AI中文摘要:

大语言模型(LLM)可以直接根据社交媒体帖子评定抑郁严重程度,或标记帖子所显示的临床标准,并让代码将计数转化为标签。后者更易于审计,因为临床医生可以检查每个标记的标准。我们在两个Reddit语料库上使用三个LLM(从9B到前沿规模)和两份问卷(PHQ-9、BDI-II)比较了这些方法,并使用二次加权kappa衡量一致性。对于两个前沿模型,仅当其决策阈值在标记数据上拟合时,标准提取在一个语料库上的得分才高于思维链。无论是否在同一标签上重新校准思维链,两个模型的增益均不显著。在根据PHQ-9标准先验固定阈值的情况下,提取在两个语料库上均无增益,即使模型在每篇帖子中标记超过两个标准。9B模型在来自抑郁社区的一个语料库上表现不同。它直接将大多数帖子标记为严重,无论是直接提示还是使用思维链,而先验规则在没有标签的情况下优于两者。在相同标签上重新校准思维链后,没有显著差距,这与校准效应一致。然而,更高的序数一致性并不能确保对严重病例的更好检测。PHQ-9标准提取漏掉了大多数严重帖子,并且从直接提示转向思维链再到提取,在几乎所有比较中增加了漏检。在主要语料库(一个重新标记的压力数据集)上,使用该数据集自身特征(包括文本中的单词计数)的模型,在先验规则下与前沿标准提取没有显著差异。

英文摘要:

A large language model (LLM) can rate depression severity directly from a social media post or mark which clinical criteria the post shows and let code turn the count into a label. The latter is easier to audit because a clinician can check each marked criterion. We compare these approaches on two Reddit corpora using three LLMs (from 9B to frontier scale) and two questionnaires (PHQ-9, BDI-II), and measure agreement with quadratic weighted kappa. For the two frontier models, criteria extraction scores above chain-of-thought on one corpus only when its decision thresholds are fitted on labeled data. Neither model's gain is significant, with or without recalibrating chain-of-thought on the same labels. With thresholds fixed a priori from PHQ-9's criteria, extraction shows no gain on either corpus, even where models mark over two criteria per post. The 9B model behaves differently on a corpus from depression communities. It labels most posts severe, whether prompted directly or with chain-of-thought, while the a priori rule beats both without labels. After chain-of-thought is recalibrated on the same labels, no significant gap remains, consistent with a calibration effect. Yet higher ordinal agreement does not ensure better detection of severe cases. PHQ-9 criteria extraction misses most severe posts, and moving from direct prompting to chain-of-thought and then to extraction increases misses in nearly all comparisons. On the primary corpus, a relabeled stress dataset, a model using that dataset's own features, including word counts from the text, is not significantly different from frontier criteria extraction under the a priori rule.

补充信息

↑