HealthBench-Psych:OpenAI的HealthBench的心理健康子集
HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench
查看机构详情
- Beth Israel Deaconess Medical Center(贝斯以色列女执事医疗中心)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
该研究针对现有通用健康基准无法单独评估心理健康性能的问题,构建了HealthBench-Psych心理健康子集,评估20个前沿及开源模型的心理健康相关表现,发现模型排名一致性高并发布了相关资源。
中文摘要 AI 辅助
通用健康基准越来越成为衡量大语言模型(LLM)医疗性能的依据,但它们并不总能按临床专科区分,导致难以单独评估特定领域的性能。心理健康是备受关注的公共卫生问题,数百万人向LLM寻求心理支持,而现有的大多数评估都是定制的学术基准,难以集成到开发者工作流中。我们推出HealthBench-Psych和HealthBench-Psych-Hard,通过透明的LLM应用规则,从HealthBench的5000条医生评分对话中筛选出与心理健康相关的内容,随后通过两轮隐藏已知排除对照的盲法临床医生审核验证该子集,最终得到610条对话(占语料库的12.2%)。在由跨厂商的3名LLM评判者组成的小组下评估20个前沿及开源模型后,我们发现前沿模型集群表现出统计显著性的平局,2个模型存在可测量的弃权(不执行)行为,且不同评判者的排名近乎一致(τ≥0.92)。我们将该子集、处理流程、模型响应、评分及分析代码作为可复用资源发布。
英文摘要
General-purpose health benchmarks increasingly anchor claims about LLM medical performance, but they are not always resolved by clinical specialty, making domain-specific performance hard to isolate. Mental health is of acute public-health concern as millions of people turn to LLMs for psychological support, and most existing evaluations are bespoke academic benchmarks that are difficult to integrate into developer workflows. We introduce HealthBench-Psych and HealthBench-Psych-Hard. We screened HealthBench's 5,000 physician-rubric conversations for mental-health relevance with a transparent LLM-applied rubric, then validated the subset through two rounds of blinded clinician review with concealed known-exclude controls, yielding 610 conversations (12.2% of the corpus). Evaluating 20 frontier and open models under a cross-vendor panel of three LLM judges, we find a statistically tied frontier cluster, measurable refusal behavior in two models, and near-identical rankings across judges ($τ\ge 0.92$). We release the subset, pipeline, model responses, grades, and analysis code as a reusable resource.