发表机构
Stanford University School of Medicine; Wu Tsai Neurosciences Institute, Stanford University; Kaiser Permanente; BigCommerce; Veterans Affairs Palo Alto Healthcare System; Sierra Pacific Mental Illness, Research, Education, and Clinical Center (MIRECC)(斯坦福大学医学院; 斯坦福大学吴蔡神经科学研究所; 凯撒医疗机构; BigCommerce公司; 帕洛阿尔托退伍军人事务医疗系统; 塞拉太平洋精神疾病研究、教育和临床中心(MIRECC))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过机制可解释性技术分析Gemma-3-27B-PT的残差流,构建与临床医生判断一致的抑郁症状向量,其第21层的抑郁向量可区分抑郁与非抑郁文本,为可解释抑郁评估工具提供机制基础。
AI 中文摘要
抑郁症患者表现出多样化的症状特征,但临床实践通常将这种多样性简化为单一的严重程度评分。大型语言模型(LLMs)有可能从患者的言语中捕捉到各种症状及其严重程度,但LLMs内部如何表征抑郁症状仍知之甚少,这限制了临床信任。为了检验模型内部激活是否与临床医生的判断一致,我们使用机制可解释性技术分析了Gemma-3-27B-PT的残差流,记录了来自经过验证的临床工具的症状描述的激活情况,发现跨多个距离指标,症状组在第21层的几何分离程度最高。随后,我们使用语义投影将保留的自然文本投影到由这些工具构建的症状向量上,得到的每个症状系数在情绪、躯体和自杀倾向轴上保留了临床医生标注的排名顺序。此外,第21层的单个抑郁向量可将保留的抑郁文本与非抑郁文本区分开来(AUC=0.789),可作为情绪效价门控,将症状投影限制在抑郁言语上。这些结果揭示了一种去相关的、与临床医生判断一致的症状信号,可直接从内部激活中读取,为可解释的抑郁评估工具提供了机制基础。
英文摘要
Patients with depression present with diverse symptom profiles, yet clinical practice routinely reduces this variation to a single severity score. Large language models (LLMs) can potentially capture various symptoms and their severity from patient speech. However, how depressive symptoms are represented inside LLMs remains poorly understood, limiting clinical trust. To examine whether internal model activations match clinician judgment, we analyzed the residual stream of Gemma-3-27B-PT using mechanistic interpretability techniques. Recording activations across symptom descriptions drawn from validated clinical instruments, we found that symptom groups geometrically separated the most at layer 21 across multiple distance metrics. Using Semantic Projection, we then projected held-out naturalistic text onto Symptom Vectors constructed from these instruments. The resulting per-symptom coefficients preserved clinician-annotated rank ordering across mood, somatic, and suicidality axes. Furthermore, a single depression vector in Layer 21 separates held-out depressive from non-depressive text (AUC = 0.789), which can be used as an emotional valence gate that restricts symptom projection to depressive speech. These results reveal a decorrelated, clinician-aligned symptom signal readable directly from internal activations, offering a mechanistic foundation for interpretable depression-assessment tools.
Comments26 pages, 6 figures