arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

生成式人工智能心理健康支持的风险治理:一种多轮安全架构

Risk Governance for Generative AI Mental Health Support: A Multi-Turn Safety Architecture

Anabela C. Areias, Catarina Botelho, António Farinhas, Areti Vassilopoulos, Dora Janela, Xin Tong, Nuno M. Guerreiro, Maya D'Eon, Fabíola Costa, Ricardo Rei

arXiv 2607.22692首次发表:更新:

AI 中文总结

研究针对大语言模型用于心理健康支持时缺乏风险治理机制的问题,开发了一种多轮安全架构,结合多种方法用于心理健康交互,经测试在多方面表现良好,能改善模型管理风险方式,提供安全部署框架。

AI 中文摘要

大语言模型(LLMs)越来越多地用于情感支持,但缺乏安全管理不断演变的心理健康风险的机制。现有安全方法主要是检测风险,很少塑造模型在对话风险展开时的响应方式。我们开发了一种与模型无关的安全治理架构,结合上下文风险检测、基于推理的验证和协议引导的响应生成,用于多轮心理健康交互。以真实世界心理健康叙述为基础的合成对话用于评估该架构的性能,在GPT-5-chat和Qwen3.5-27B上进行测试,实现了高风险检测性能(特异性:0.85(95%CI:0.78;0.91),敏感性:0.92(95%CI:0.88;0.95)),并将临床医生首选的升级响应提高了25.6 - 59.2个百分点,同时保持融洽关系和联系。性能在对话长度上保持稳定,并在专有和开源模型中普遍适用。这些发现表明,基于临床的安全治理可以超越风险检测,改善LLMs管理不断演变的心理健康风险的方式,为跨模型更安全的部署提供可扩展框架。

英文摘要

Large language models (LLMs) are increasingly used for emotional support despite lacking mechanisms to safely govern evolving mental health risk. Existing safety approaches primarily detect risk but rarely shape how models respond as conversational risk unfolds. We developed a model-agnostic safety governance architecture that combines contextual risk detection, reasoning-based verification, and protocol-guided response generation for multi-turn mental health interactions. Synthetic conversations grounded in real-world mental health narratives were used to evaluate the architecture's performance, tested with GPT-5-chat and Qwen3.5-27B, achieving high risk detection performance (specificity: 0.85 (95\%CI: 0.78;0.91), sensitivity: 0.92 (95\%CI: 0.88;0.95)) and increasing clinician-preferred escalation responses by 25.6--59.2pp while preserving rapport and connection. Performance remained stable across conversation length and generalized across both proprietary and open-source models. These findings demonstrate that clinically-grounded safety governance can extend beyond risk detection to improve how LLMs manage evolving mental health risk, providing a scalable framework for safer deployment across models.

DOI:10.21203/rs.3.rs-10279940/v1

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑