arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20345cs.CLcs.AIcs.CY

当词汇理解失效于临床推理:评估面向阿尔法世代(Gen Alpha,2010-2024年出生)的治疗机器人的安全风险

When Vocabulary Comprehension Fails Clinical Reasoning: Evaluating Therapy Bots' Safety Risks for Generation Alpha

Manisha Mehta, Virendra Mehta

首次发表
浏览论文内容

中文总结 AI 辅助

本研究构建两个基准评估Claude、GPT-4o等LLM的治疗机器人安全风险,发现其存在10-14个百分点的词汇理解与临床风险校准差距,识别六种失败模式,提出四项监管与技术改进建议。

中文摘要 AI 辅助

对话式AI系统已成为阿尔法世代(Gen Alpha,2010-2024年出生)的非正式心理健康支持资源,美国13.1%的青少年(约540万人)使用生成式AI获取心理健康建议。尽管这些系统(从治疗应用到通用聊天机器人)基于在大量心理学文献上训练的大语言模型(LLM)构建,但它们对以夸张语言、反讽式积极表达、快速语义漂移和语境多义性为特征的青少年沟通模式的安全性仍未得到验证。在多起与AI聊天机器人互动相关的青少年死亡事件后,系统性评估至关重要。本研究提出两个基准:(1)64个经母语使用者(组内相关系数ICC=0.72)和临床医生(科恩kappa系数=0.78)验证的阿尔法世代心理健康表达;(2)75个多轮对话(共780轮),包含标准版本与阿尔法世代版本的配对。在对治疗应用和通用聊天机器人底层的LLM架构(Claude、GPT-4o、Llama-3.1)的评估中,模型对词汇的理解率为76%-82%,但临床风险校准正确率仅为64%-72%,形成10-14个百分点(pp)的词汇理解差距(p<.001,d>0.48),而人类治疗师的该差距仅为3pp(p=.22)。该差距在架构上具有一致性,且随歧义性增大而扩大(7pp→18pp)。研究识别出六种失败模式:讽刺掩盖(29pp)、最小化态度接受(43pp)、非正式风格偏差(24pp)、风险分层歧义(19pp)、语义漂移(19pp)、语境依赖型暴力(7pp)。这些模式会产生叠加效应;存在三种及以上模式时,风险漏检率达94%。轻量缓解措施无效;仅重度支架式干预可达到人类表现水平(成本为6.4倍)。鉴于34%的基线漏检率对应每年约146880起漏检危机,本研究建议强制实施人在回路(human-in-the-loop)架构、每季度开展针对青少年的验证、公开透明的性能披露,以及针对面向青少年的心理健康AI的监管框架。

英文摘要

Conversational AI systems have become informal mental health support resources for Generation Alpha (Gen Alpha, born 2010-2024), with 13.1% of U.S. adolescents (5.4 million) using generative AI for mental health advice. While these systems, from therapy apps to general chatbots, rely on large language models trained on extensive psychological literature, their safety for youth communication patterns characterized by hyperbolic language, ironic positivity, rapid semantic drift, and contextual polysemy remains unvalidated. Following multiple adolescent deaths linked to AI chatbot interactions, systematic evaluation is critical. We present two benchmarks: (1) 64 Gen Alpha mental health expressions validated by native speakers (ICC=0.72) and clinicians (kappa=0.78); (2) 75 multi-turn conversations (780 turns) with paired Standard/Gen Alpha versions. Across evaluations of LLM architectures underlying therapy apps and general chatbots - Claude, GPT-4o, Llama-3.1 - models understand 76-82% of vocabulary but correctly calibrate only 64-72% of clinical risk, creating a 10-14 percentage point (pp) vocabulary-comprehension gap (p<.001, d>0.48) absent in human therapists (3pp, p=.22). The gap is architecturally consistent and widens with ambiguity (7pp -> 18pp). We identify six failure patterns: sarcasm masking (29pp), minimization acceptance (43pp), informal style bias (24pp), risk-stratified ambiguity (19pp), semantic drift (19pp), context-dependent violence (7pp). Patterns compound; three or more yield 94% miss rates. Lightweight mitigations fail; only heavy scaffolding achieves human performance (6.4x cost). With 34% baseline miss rate yielding 146,880 estimated annual missed crises, we recommend mandatory human-in-the-loop architectures, quarterly youth-specific validation, transparent performance disclosure, and regulatory frameworks for youth-facing mental health AI.

发表机构

  • Lynbrook High School(林布鲁克高中)
  • University of Trento(特伦托大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑