arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20684cs.CL

HerHealthEval:评估女性健康沟通的多语言与语域敏感理解

HerHealthEval: Evaluating Multilingual and Register-Sensitive Understanding of Women's Health Communication

Hassan Saeed Hassan Albattra, Mazen Mohammed Bahgat, Rahatara Ferdousi, Hana Essam Sayed Ahmed Amrya, Mariam Mousa

首次发表
浏览论文内容

中文总结 AI 辅助

针对多语言女性健康沟通评估忽视语域变化的问题,提出HerHealthEval框架,通过六种沟通形式测试模型理解,发现适配标签来源影响欠分流风险,强调评估需关注语域与不确定性。

中文摘要 AI 辅助

大型语言模型越来越多地用于医疗保健沟通,然而大多数评估侧重于响应质量,同时假设用户的关切已被正确解读。我们引入了HerHealthEval,一个用于女性健康沟通多语言理解的可控评估框架。对于每个临床案例,HerHealthEval提供英语、法语和现代标准阿拉伯语的匹配版本,使用六种沟通形式:规范式、临床式、非专业式、间接或委婉式、情绪关切式以及故意信息不足式。前五种表达相同的潜在关切并保留相同的临床信息,而信息不足式则有意省略相关细节,以测试模型是否认识到需要澄清。我们评估了一个多语言指令模型及QLoRA适配变体在关切分类、风险校准、澄清行为、解析合规性和跨形式一致性方面的表现。结果显示,总体准确率和一致性可能掩盖安全相关的失败。一个多语言适配模型在语言不对称风险监督下,在法语和阿拉伯语中达到0.994的欠分流率。使用源自源语言、语言不变的风险标签进行受控重新适配,分别将欠分流率降至0.572和0.558。这些发现表明,稳健的多语言医疗保健评估需要明确测试语域变化、不确定性处理以及适配标签的来源和不变性。

英文摘要

Large language models are increasingly used in healthcare communication, yet most evaluations emphasize response quality while assuming that the user's concern has been interpreted correctly. We introduce HerHealthEval, a controlled evaluation framework for multilingual understanding of women's-health communication. For each clinical case, HerHealthEval provides matched versions in English, French, and Modern Standard Arabic using six communicative forms: canonical, clinical, layperson, indirect or hedged, emotionally concerned, and deliberately under-specified. The first five express the same underlying concern and retain the same clinical information, whereas the under-specified form intentionally omits relevant details to test whether the model recognizes that clarification is needed. We evaluate a multilingual instruction model and QLoRA-adapted variants on concern classification, risk calibration, clarification behavior, parse compliance, and cross-form consistency. Results reveal that aggregate accuracy and consistency can conceal safety-relevant failures. A multilingual adaptation model reaches 0.994 under-triage in French and Arabic under language-asymmetric risk supervision. A controlled re-adaptation using source-derived, language-invariant risk labels reduces under-triage to 0.572 and 0.558, respectively. These findings show that robust multilingual healthcare evaluation requires explicit testing of register variation, uncertainty handling, and the provenance and invariance of adaptation labels.

补充信息

↑