arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

K-Bench:一个临床校准的基准,用于评估高风险心理健康对话中的大型语言模型

K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations

Laura M. Vowels, Matthew J. Vowels, Shivali Sharma, Apoorv Jha, Rehnuma Choudhury, Wasseem El Sarraj, Rachel Francois-Walcott, Aruba Hussain, Sarah Ingram, Angela Loulopoulou, Adva Segal, Elena Volkova

arXiv 2609.15855首次发表:更新:

发表机构

University of Roehampton; Kivira Health; University of Hertfordshire; University of Surrey; University of Bedfordshire; Tavistock Relationships; InsideOut(罗汉普顿大学; Kivira Health; 赫特福德大学; 萨里大学; 贝德福德大学; 塔维斯托克关系中心; InsideOut)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

K-Bench是一个临床校准的基准,通过200个多轮高风险心理健康情景评估125种LLM配置,发现领先模型综合风险评分超95,且治疗性提示仅对较弱模型有效。

AI 中文摘要

人们越来越多地使用大型语言模型(LLMs)来提供心理健康支持,然而它们在不断演变的高风险对话中的安全性仍然缺乏充分描述。我们开发了K-Bench,一个临床医生校准的、受保护的基准,评估了来自14个提供商的33个基础模型的125种模型配置,这些配置在一个包含200个多轮情景的固定队列上进行测试,情景涉及自杀、自残、家庭暴力、物质滥用以及无风险表现。合成的患者对话与真实的人机对话在分布上显示出显著重叠。一个冻结的GPT-4o评审在来自151个临床医生评分的转录本的6,751个符合条件的项目比较中,与临床医生共识达到了94.2%的精确一致。领先的模型将强大的支持性对话与综合风险评分超过95相结合,而风险探索则暴露出较低性能配置之间的显著差异。治疗性提示在较弱的模型中产生了特定配置的改进,而提高推理能力并未带来平均改善。K-Bench结合了更广泛的临床覆盖和配置规模的比较,并提供了一个持续更新的公共排行榜,其操作测试材料受到保护,防止直接优化。排行榜可在该http URL获取。

英文摘要

People increasingly use large language models (LLMs) for mental health support, yet their safety in evolving, high-risk conversations remains poorly characterised. We developed K-Bench, a clinician-calibrated, protected benchmark evaluating 125 model configurations representing 33 base models from 14 providers across a fixed cohort of 200 multi-turn vignettes involving suicide, self-harm, domestic violence, substance misuse, and no-risk presentations. Synthetic patient conversations showed substantial distributional overlap with real human-AI conversations. A frozen GPT-4o judge achieved 94.2% exact agreement with clinician consensus across 6,751 eligible item comparisons from 151 clinician-rated transcripts. Leading models combined strong supportive conversation with combined-risk scores above 95, whereas risk exploration exposed substantial variation among lower-performing configurations. Therapeutic prompting produced configuration-specific gains concentrated among weaker models, while elevated reasoning produced no average improvement. K-Bench combines broader clinical coverage and configuration-scale comparison with a continuously updated public leaderboard whose operational test materials are protected from direct optimisation. The leaderboard is available at www.k-bench.ai.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑