发表机构
INRIA; Carl von Ossietzky University of Oldenburg; German Research Center for Artificial Intelligence; Max Planck Institute for Intelligent Systems(法国国家信息与自动化研究所; 奥尔登堡大学; 德国人工智能研究中心; 马克斯·普朗克智能系统研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究结合访谈者判断与Ridge、BiLSTM等自动语言分析模型,在107次精神科临床对话数据上实现了更准确的患者体验评估,证明二者可互补提升评估效果。
AI 中文摘要
理解精神科患者对临床对话的主观体验,对反馈和与联盟相关的过程监测十分重要。访谈者会在诊疗结束后对患者体验形成判断,但这些判断并不总是与患者的自我报告一致。已有研究提出了从对话预测感知互动质量的自动方法,但目前尚不清楚这类方法能否补充人类判断,而非仅仅复制人类判断。为解决这一空白,我们评估了一种临床医生支持框架,该框架将诊疗结束后访谈者的评分与基于语言的自动预测相结合,以估计自由临床访谈中患者报告的互动质量。我们在多种标准模型类型上评估了这种整合,包括Ridge、SVR、MLP、GRU和BiLSTM,所有模型均基于从精神科患者与访谈者的107次自由对话的二元转录本中提取的句子嵌入进行训练。我们的结果显示,通过简单平均将访谈者判断与模型预测相结合,可获得最强的整体性能。仅访谈者的基准模型达到了0.365的皮尔逊相关系数;在全自动化模型中,Ridge达到了最强的皮尔逊相关系数(r=0.286),BiLSTM达到了r=0.270;最强结果由BiLSTM与访谈者判断整合得到,其皮尔逊相关系数为0.403。我们的研究结果表明,自动语言分析与访谈者判断捕捉了患者体验的互补方面,二者的结合比单独使用任一来源都能更准确地近似患者的自我报告。
英文摘要
Understanding how psychiatric patients subjectively experienced a clinical conversation is important for feedback and alliance-related process monitoring. While interviewers form post-session judgments about patient experience, these judgments do not always match patients' self-reports. Automatic approaches for predicting perceived interaction quality from conversation have been proposed, but it remains unclear whether such approaches can complement human judgment rather than simply replicate it. To address this gap, we evaluate a clinician-support framework in which post-session interviewer ratings are combined with automatic language-based predictions to estimate patient-reported interaction quality in free clinical interviews. We assess this integration across multiple standard model types, including Ridge, SVR, MLP, GRU, and BiLSTM, all trained on sentence embeddings extracted from dyadic transcripts of 107 free conversations between psychiatric patients and interviewers. Our results show that combining interviewer judgments with model predictions through simple averaging yields the strongest overall performance. The interviewer-only baseline reached a Pearson correlation of 0.365. Among fully automatic models, Ridge achieved the strongest Pearson correlation (r = 0.286), while BiLSTM achieved r = 0.270. The strongest result was obtained by BiLSTM interviewer integration (r = 0.403). Our findings suggest that automatic language analysis and interviewer judgment capture complementary aspects of patient experience and that their combination provides a more accurate approximation of the patient's own report than either source alone.
CommentsAccepted to ACII 2026 as an oral presentation