发表机构
University of North Carolina at Chapel Hill; University of Auckland; Purdue University; East River Counseling(北卡罗来纳大学教堂山分校; 奥克兰大学; 普渡大学; 东河咨询机构)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出临床医生参与的基准,评估LLMs能否从多模态证据生成B-HiTOP条目画像,发现两阶段预测可提升部分证据的兼容性,但语义抽象会成为间接行为传感的信息瓶颈。
AI 中文摘要
移动与可穿戴传感技术支持对行为的纵向观测,但将这些信号转化为有意义的心理健康构念仍存在困难。我们提出一种有临床医生参与的基准,用于评估大型语言模型(LLMs)能否从被动传感、生态瞬时评估(EMA)及问卷证据中生成基于证据的简短精神病理学层次分类(B-HiTOP)条目画像。利用纵向行为建模泛化(GLOBEM)数据集,我们构建了14592个参与者-日实例,并将多模态证据与五个维度的29个B-HiTOP条目对齐。由于GLOBEM数据集缺乏B-HiTOP响应,我们评估证据兼容性(C)而非诊断准确性,当证据不足以进行条目级评分时,将实质性预测与弃权(不执行)区分开。两阶段预测可提高EMA和问卷证据的C值,但会降低被动传感及组合证据下的C值,并在各模型、维度及证据设置中产生更保守的分数分布。总体而言,语义抽象有助于组织异构的自我报告证据,但会成为间接行为传感信号的信息瓶颈。
英文摘要
Mobile and wearable sensing enables longitudinal observation of behavior, yet translating these signals into meaningful mental health constructs remains difficult. We introduce a clinician-in-the-loop benchmark for evaluating whether large language models (LLMs) can generate evidence-grounded Brief Hierarchical Taxonomy of Psychopathology (B-HiTOP) item profiles from passive sensing, ecological momentary assessment (EMA), and questionnaire evidence. Using the Generalization of Longitudinal Behavior Modeling (GLOBEM) dataset, we construct 14,592 participant-day instances and align multimodal evidence to 29 B-HiTOP items across five spectra. Since GLOBEM lacks B-HiTOP responses, we evaluate evidence compatibility (C) rather than diagnostic accuracy, separating substantive predictions from abstentions when evidence is insufficient for item-level scoring. Two-stage prediction improves C for EMA and questionnaire evidence, but reduces C under passive sensing and combined evidence and produces more conservative score distributions across models, spectra, and evidence settings. Overall, semantic abstraction helps organize heterogeneous self-report evidence while becoming an information bottleneck for indirect behavioral sensing signals.