arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CARE-MH:迈向心理健康语言模型的统一、可重现和可比评估

CARE-MH: Towards Unified, Reproducible, and Comparable Evaluation of Mental Health LLMs

Asher Sprigler, Yixue Zhao, Yi Ding

arXiv 2607.24754首次发表:更新:

发表机构

Purdue University; Yixue Research Institute(普渡大学; 益学研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对心理健康语言模型评估难重现、难比较的问题,提出CARE-MH统一框架,通过重现分析现有基准,发现可重复性取决于模型稳定性,跨基准分歧源于指标定义差异,强调了标准化评估配置和共享指标定义的必要性。

AI 中文摘要

大语言模型(LLMs)越来越多地用于提供心理健康支持,这需要对安全性、同理心和治疗适宜性进行可靠评估。然而,由于评估设计和指标定义不一致,现有的心理健康基准难以重现和比较。我们提出了CARE-MH,这是一个用于心理健康LLMs可比和可重现评估的统一框架。使用CARE-MH,我们重现并分析了最新基准,发现可重复性很大程度上取决于模型稳定性,跨基准分歧主要源于指标定义的差异。我们的发现凸显了未来心理健康LLM基准需要标准化评估配置和共享指标定义。

英文摘要

Large language models (LLMs) are increasingly used to provide mental health support, requiring reliable evaluation of safety, empathy, and therapeutic appropriateness. However, existing mental health benchmarks are difficult to reproduce and compare due to inconsistent evaluation designs and metric definitions. We present CARE-MH, a unified framework for comparable and reproducible evaluation of mental health LLMs. Using CARE-MH, we reproduce and analyze state-of-the-art benchmarks, revealing that reproducibility depends strongly on model stability and that cross-benchmark disagreement primarily arises from differences in metric definitions. Our findings highlight the need for standardized evaluation configurations and shared metric definitions for future mental health LLM benchmarks.

Comments39 pages, 22 figures, 22 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑