评估大语言模型对海地克里奥尔语的文化意识
Evaluating Cultural Awareness of LLMs for Haitian Creole
- Telecom Paris, Institut Polytechnique de Paris(巴黎电信学院,巴黎综合理工学院)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本研究首次系统评估大语言模型在海地克里奥尔语中的文化意识,发现其与法语存在明显差距,且故事生成中即使正面刻画也常编码刻板印象,代码和基准已公开。
中文摘要 AI 辅助
大语言模型(LLMs)在高资源语言与低资源语言之间表现出显著的性能差异。除了较低的任务性能外,它们往往无法捕捉代表性不足社区的文化规范和价值观。在这项工作中,我们首次对LLMs在海地克里奥尔语中的文化意识进行了系统性评估,该语言有数百万人使用,但在数字资源中严重代表性不足。我们通过四个互补维度——特异性、偏见、多样性和变化——评估文化意识,使用由母语者在文本填充环境中策划的文化显著提示基准。我们的结果揭示了海地克里奥尔语与高资源法语之间文化意识的明显差距,海地语的表现跨领域更不均匀,且更受法语语言干扰的影响。故事生成进一步揭示了通过艰辛和韧性反复描绘海地角色的现象,表明即使是正面的刻画也可能编码刻板印象叙事。我们的代码、基准和评估框架已公开提供。
英文摘要
Large language models (LLMs) exhibit substantial performance disparities between high- and low-resource languages. Beyond lower task performance, they often fail to capture the cultural norms and values of underrepresented communities. In this work, we present the first systematic evaluation of cultural awareness in LLMs for Haitian Creole, a language spoken by millions but severely underrepresented in digital resources. We assess cultural awareness along four complementary dimensions---specificity, bias, diversity, and variation---using a benchmark of culturally salient prompts curated by native speakers in a text infilling setting. Our results reveal a clear gap between cultural awareness in Haitian Creole and higher-resource French, with Haitian performance being more uneven across domains and more affected by French linguistic interference. Story generation further reveals recurring portrayals of Haitian characters through hardship and resilience, showing that even positive characterizations can encode stereotypical narratives. Our code, benchmark, and evaluation framework are publicly available.