面向句子级抑郁症状识别的候选生成与定义引导验证
Candidate Generation and Definition-Guided Verification for Sentence-Level Depression Symptom Recognition
另 1 家 · 查看机构详情
- Institute for Systems and Robotics (ISR)(系统与机器人研究所)
- Instituto Superior Técnico (IST)(高等技术学院)
- University of Lisbon(里斯本大学)
- Hospital Beatriz Ângelo(Beatriz Ângelo医院)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究针对句子级抑郁症状识别的挑战,提出两阶段框架,经对比微调的句子编码器生成症状候选,微调后的LLM结合定义验证,在多基线对比中取得最优准确率与F1分数,支持症状识别的两阶段分解思路。
中文摘要 AI 辅助
句子级抑郁症状识别颇具挑战性,因为相似表达在症状相关性上可能存在差异,且语言模型的推理缺乏诊断定义的支撑。本研究提出一个两阶段框架,将症状候选生成与基于定义的验证分离开来:经对比微调的句子编码器为每个句子生成一个症状候选,微调后的语言模型则利用该句子、其上下文以及候选特定的诊断定义,验证该候选是否存在,在给出答案前会依据定义核查自身判断。与编码器、基于推理的、医学及通用大型语言模型(LLM)基线,以及匹配的单阶段监督分类器相比,所提流程在所有方法中取得了最佳的准确率与F1分数,且其依据与专家编写的注释相符。初步临床审计显示其与诊断定义存在中等程度的一致性,解释质量高度依赖预测的正确性。研究结果支持将症状识别分解为候选生成与基于定义的验证,不过在稀有类别上的性能仍有限。
英文摘要
Sentence-level recognition of depression symptoms is challenging because similar expressions can differ in symptom relevance, and language-model inference is insufficiently grounded in diagnostic definitions. This study proposes a two-stage framework separating symptom-candidate generation from definition-grounded verification. A contrastively fine-tuned sentence encoder generates a symptom candidate per sentence, and a fine-tuned language model verifies whether the candidate is present or absent using the sentence, its context, and a candidate-specific diagnostic definition, checking its judgment against that definition before answering. Evaluated against encoder, inference-based, medical, and general LLM baselines and a matched single-stage supervised classifier, the proposed pipeline attains the best accuracy and F1 scores of all methods, with rationales matching expert-authored annotations. A preliminary clinical audit indicates moderate alignment with diagnostic definitions, with explanation quality strongly dependent on prediction correctness. The results support decomposing symptom recognition into candidate generation and definition-grounded verification, though performance remains limited for rare categories.