探测支持语音的大语言模型(LLM)在心理健康对话中由温暖感介导的伤害
Probing Warmth-Mediated Harm in Speech-Enabled LLMs for Mental-Health Conversations
- Microsoft(微软)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对支持语音的LLM,基于WHO指南设计7轮心理健康对话探测,发现音频模态下模型存在温暖感缺失的特定模式,强调需结合音频文本评估模型,发布了相关评估资源。
AI中文摘要:
音频大语言模型(LLM)基准测试衡量的是理解能力和对话质量,而非当脆弱用户披露心理健康问题时,支持语音的模型是否会以关系温暖感做出回应。我们引入了一个基于世界卫生组织(WHO)心理健康临床指南的7轮脚本式披露探测,每个脚本都在同一模型(Azure OpenAI gpt-realtime)上以音频和仅文本两种条件运行,并对生成的语音进行声学韵律分析。在532个回应中,我们发现了两种仅文本转录评估会遗漏的音频特有模式:在引导轮次,模型的声音变得更短、更快、音高更低且更安静,而非更温暖(7项声学特征中有5项的p值小于0.001);关系接纳的模态差距在整体上较小,但集中在风险最高的自我伤害/自杀脚本中。一项由两名评分者参与的听众研究证实,感知到的温暖感集中在特定轮次和丧亲披露上。这些模式共同表明,在心理健康情境中审计支持语音的模型需要评估用户遇到的音频与文本结合的体验,而非仅转录内容。我们发布了该协议、评分流程和脚本,作为评估心理健康情境中支持语音的模型的起点。
英文摘要:
Audio LLM benchmarks measure understanding and dialogue quality, not whether speech-enabled models respond with relational warmth when a vulnerable user discloses a mental-health concern. We introduce a 7-turn scripted-disclosure probe grounded in WHO mental-health clinical guidelines, with each script run on the same model (Azure OpenAI gpt-realtime) in both audio and text-only conditions, and acoustic-prosody analysis of the generated speech. Across 532 responses we identify two audio-specific patterns transcript-only evaluation would miss: at the elicitation turn the model's voice gets shorter, faster, lower-pitched, and quieter rather than warmer (p < .001 for five of seven acoustic features), and the modality gap on relational acceptance, small in aggregate, concentrates in the highest-stakes self-harm/suicide scripts. A two-rater listener study corroborates that perceived warmth is concentrated at specific turns and on bereavement disclosures. Together these patterns indicate that auditing speech-enabled models in mental-health contexts requires evaluating the combined audio-and-text experience the user encounters, not the transcript in isolation. We release the protocol, scoring pipeline, and scripts as a starting point for evaluating speech-enabled models in mental-health contexts.