发表机构
Microsoft Research Africa; Swansea University; Cornell University; Microsoft Research Cambridge(微软研究院非洲分部; 斯旺西大学; 康奈尔大学; 微软研究院剑桥分部)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文以五个非洲社区为对象,通过混合方法评估AI生成多模态故事的文化对齐情况,构建了相关分类体系,并发现多模态LLM裁判需经社区校准才能可靠评估文化对齐。
AI 中文摘要
本文研究AI生成的多模态故事与所代表社区的生活实践、人际关系、语言、价值观及视觉期望的契合程度。我们在五个非洲社区开展基于社区的混合方法评估,涉及19名文化代表,结合定量标注与定性焦点小组讨论。研究发现,文化对齐不仅取决于可识别的文化标记,更取决于这些标记如何适配社会、语言、程序及视觉语境。基于评估结果,我们构建了文化对齐分类体系,包含五大类文化标记及八大类常见的不对齐机制。我们还评估了五个多模态大语言模型(LLM)裁判,以检验自动化评估能否大规模近似基于社区的判断。不同社区的裁判可靠性与分数校准差异显著,没有单一裁判在所有五个场景中表现一致。这些发现推动了基于社区校准的评估流程,即通过社区判断验证自动化裁判,确定其可信应用场景及需人工审核的环节。
英文摘要
In this paper, we examine how well AI-generated multimodal stories align with the lived practices, relationships, language, values, and visual expectations of the communities they represent. We conduct a community-grounded mixed-methods evaluation with 19 culture representatives across five African communities, combining quantitative annotations with qualitative focus group discussions. We find that cultural alignment depends not simply on recognizable cultural markers, but on how those markers fit social, linguistic, procedural, and visual context. From these evaluations, we develop a taxonomy of cultural alignment comprising five broader cultural marker categories and eight recurring mechanisms of misalignment. We additionally evaluate five multimodal LLM judges to examine whether automated evaluation can approximate community-grounded judgments at scale. Judge reliability and score calibration vary substantially across communities, with no single judge performing consistently across all five settings. These findings motivate community-calibrated evaluation pipelines in which automated judges are validated against community judgments to determine where they can be trusted and where human review remains necessary.
CommentsAccepted to Findings of EMNLP 2026