HerHealthEval:评估女性健康沟通的多语言与语域敏感理解
HerHealthEval: Evaluating Multilingual and Register-Sensitive Understanding of Women's Health Communication
浏览论文内容
中文总结 AI 辅助
针对多语言女性健康沟通评估忽视语域变化的问题,提出HerHealthEval框架,通过六种沟通形式测试模型理解,发现适配标签来源影响欠分流风险,强调评估需关注语域与不确定性。
中文摘要 AI 辅助
大型语言模型越来越多地用于医疗保健沟通,然而大多数评估侧重于响应质量,同时假设用户的关切已被正确解读。我们引入了HerHealthEval,一个用于女性健康沟通多语言理解的可控评估框架。对于每个临床案例,HerHealthEval提供英语、法语和现代标准阿拉伯语的匹配版本,使用六种沟通形式:规范式、临床式、非专业式、间接或委婉式、情绪关切式以及故意信息不足式。前五种表达相同的潜在关切并保留相同的临床信息,而信息不足式则有意省略相关细节,以测试模型是否认识到需要澄清。我们评估了一个多语言指令模型及QLoRA适配变体在关切分类、风险校准、澄清行为、解析合规性和跨形式一致性方面的表现。结果显示,总体准确率和一致性可能掩盖安全相关的失败。一个多语言适配模型在语言不对称风险监督下,在法语和阿拉伯语中达到0.994的欠分流率。使用源自源语言、语言不变的风险标签进行受控重新适配,分别将欠分流率降至0.572和0.558。这些发现表明,稳健的多语言医疗保健评估需要明确测试语域变化、不确定性处理以及适配标签的来源和不变性。
英文摘要
Large language models are increasingly used in healthcare communication, yet most evaluations emphasize response quality while assuming that the user's concern has been interpreted correctly. We introduce HerHealthEval, a controlled evaluation framework for multilingual understanding of women's-health communication. For each clinical case, HerHealthEval provides matched versions in English, French, and Modern Standard Arabic using six communicative forms: canonical, clinical, layperson, indirect or hedged, emotionally concerned, and deliberately under-specified. The first five express the same underlying concern and retain the same clinical information, whereas the under-specified form intentionally omits relevant details to test whether the model recognizes that clarification is needed. We evaluate a multilingual instruction model and QLoRA-adapted variants on concern classification, risk calibration, clarification behavior, parse compliance, and cross-form consistency. Results reveal that aggregate accuracy and consistency can conceal safety-relevant failures. A multilingual adaptation model reaches 0.994 under-triage in French and Arabic under language-asymmetric risk supervision. A controlled re-adaptation using source-derived, language-invariant risk labels reduces under-triage to 0.572 and 0.558, respectively. These findings show that robust multilingual healthcare evaluation requires explicit testing of register variation, uncertainty handling, and the provenance and invariance of adaptation labels.