arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25287cs.HC

大语言模型能否识别并修复关系破裂?临床医生实践与大语言模型行为的比较

Can LLMs identify and repair ruptures? Comparison between clinician practices and LLM behaviors

  • University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)
  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
  • Indiana University Indianapolis(印第安纳大学印第安纳波利斯分校)

机构由 AI 辅助整理,请以论文原文为准。

Jeongah Lee, Joy Qiuyue Zhong, Drishti Goel, Violeta J. Rodriguez, Dong Whi Yoo, Koustuv Saha, Ravi Karkar

AI总结:

本研究通过情景驱动实证比较三个大语言模型与22位专家在21个心理健康对话中识别和修复关系破裂的表现,发现大语言模型在识别上更依赖显性线索且与标签一致,但在解决策略上不如专家有效,并据此提出对话代理设计需注重关系意识、节奏和人在回路支持。

AI中文摘要:

关系破裂(Ruptures)代表互动中关系对齐失效的常见但关键的时刻,这使得它们对于评估在信任和参与度最为重要的人工智能至关重要。在一项情景驱动的实证研究中,我们考察了三个大语言模型在21个心理健康对话中识别和解决关系破裂的表现,以及22位专家对这些策略的评估。在识别方面,大语言模型依赖于单轮对话中的显性语言线索,而专家则整合了跨对话的隐性、关系性和情境性信息。在解决方面,大语言模型倾向于产生更具指导性和脚本化的回应,而专家则采用过程导向的策略,如验证、开放式探索和心理教育。总体而言,大语言模型在识别方面与预定义标签的一致性更高,但在解决方面并非如此,专家对其回应的有效性评价仅为中等,且在时机、深度和情境敏感性方面存在一致的局限性。我们讨论了对心理健康对话代理设计的启示,强调关系意识、节奏和人在回路支持。

英文摘要:

Ruptures represent common albeit critical moments in interaction where relational alignment breaks down, making them essential for evaluating AI where trust and engagement matter most. In a scenario-driven empirical study, we examined the performance of three LLMs at identifying and resolving ruptures across 21 mental health conversations and 22 experts' evaluation of the strategies. For identification, LLMs relied on explicit linguistic cues within single turns whereas experts integrated implicit, relational, and contextual information across the conversation. For resolution, LLMs tended to produce more directive and scripted responses whereas experts adopted process-oriented strategies such as validation, open-ended exploration, and psychoeducation. Overall, LLMs showed higher agreement with predefined labels in identification, but not in resolution where experts rated their responses only moderately effective, with consistent limitations in timing, depth, and contextual sensitivity. We discuss implications for the design of mental health conversational agents emphasizing relational awareness, pacing, and human-in-the-loop support.

↑