arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越平均水平的文化对齐度测量:印度语境下孕产妇健康大语言模型交互评估框架

Measuring Cultural Alignment Beyond the Average: A Framework for Evaluating Maternal-Health LLM Interactions in Indian Contexts

Umaira Izhar, Gunjan Arora, Pushpendra Singh

arXiv 2610.11586首次发表:更新:

发表机构

Indraprastha Institute of Information Technology Delhi(因德拉普拉斯塔信息技术学院德里分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对印度北部语境提出MH-INDIC框架评估孕产妇健康LLM交互,发现现有LLM的文化差异感知弱于人类,且人类基础提示可提升对话的文化适配性与个体对齐度。

AI 中文摘要

现有医疗大语言模型(LLM)的评估方法主要评估事实正确性、安全性和流畅性,却难以判断生成的交互是否反映了基于文化情境的医疗推理,这一局限在孕产妇健康领域尤为关键,因为医疗决策受社会和关系规范的影响。我们提出MH-INDIC,一个针对印度北部城市及半城市语境的孕产妇健康交互文化基础评估框架,通过孕产妇健康推理的十个维度将文化行为具体化。我们对来自印度北部城市及半城市的102名孕妇和产后女性开展了含26个题项的调查,以此评估十个LLM。我们区分了人口层面的文化对齐度与个体层面的行为差异。尽管部分模型接近人类人口层面的分布,但所有被评估系统在人口统计和家庭层面的差异均显著低于人类群体,揭示了总体对齐度与个体条件敏感性之间存在差距。作为MH-INDIC的下游应用,我们使用对齐度最高的专有及开源模型,在零样本、自条件和人类基础提示下生成文化适配的孕产妇健康对话。人类基础条件下生成的对话具有更强的个体层面对齐度和更高的对话质量评分,表明测量得到的文化特征可提升生成交互的文化适配性。

英文摘要

Existing evaluation methods for healthcare LLMs primarily assess factual correctness,safety, and fluency, while providing limited insight into whether generated interactions reflect culturally situated healthcare reasoning. This limitation is particularly important in maternal health, where care decisions are shaped by social and relational norms. We introduce MH-INDIC, a culturally grounded evaluation framework for maternal-health interactions in urban and semi-urban North Indian contexts that operationalises cultural behaviour through ten dimensions of maternal-health reasoning. Using a 26-item survey administered to 102 pregnant and postpartum women from urban and semi-urban North India, we evaluate ten LLMs. We distinguish population level cultural alignment from profile-level behavioural variation. Although several models approximate the human population-level distribution, all evaluated systems exhibit substantially lower variation across demographic and household profiles than the human cohort, revealing a gap between aggregate alignment and profile-conditioned sensitivity. As a downstream application of MH-INDIC, we use the strongest-aligned proprietary and open-source models to generate culturally conditioned maternal-health dialogues under zero-shot, self-conditioned, and human-grounded prompting. Human-grounded conditioning produces stronger profile alignment and dialogue quality ratings, suggesting that measured cultural profiles can improve the cultural grounding of generated interactions

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑