大型语言模型低估并在一定程度上歪曲日常规范中的文化差异
Large language models underestimate and partly misrepresent cultural variation in everyday norms
浏览论文内容
中文总结 AI 辅助
本研究利用GSEN基准测试四个LLM,发现它们显著低估文化差异幅度且模式识别不佳,本地语言提示仅适度改善,表明LLM对日常规范的文化差异表征薄弱。
中文摘要 AI 辅助
文化的一个关键方面是社会关于日常行为的规范。大型语言模型(LLMs)在多大程度上准确地代表了此类规范中的文化差异?为了回答这个问题,我们使用了最近的全球日常规范研究(GSEN),该研究收集了90个社会中150个情景的评分,作为人类基准。我们提示GPT-5估计每个社会对每个情景的平均评分,随后在另外三个LLM中重复了该基准测试:GPT-5.4、Claude Opus 4.6和Gemini 3.1 Pro。与GSEN的估计相比,所有四个LLM都以两种方式歪曲了文化差异。首先,它们大大低估了其幅度,估计社会之间的差异平均不到其测量大小的一半。其次,对于许多情景,LLM未能很好地识别差异模式,即哪些社会认为行为不太可接受,哪些社会认为更可接受。对于引发粗俗关注的情景,尤其是涉及接吻和调情的情景,差异模式的识别效果更好。我们还发现,较发达社会的规范往往被估计得稍微更准确一些,而使用当地调查语言而非英语进行提示仅带来了准确性的适度提升。本地语言提示也减少但并未消除对社会间差异的低估。LLM对日常规范中的文化差异表征较弱且不均匀。
英文摘要
A key aspect of culture is a society's norms about everyday behavior. How accurately do large language models (LLMs) represent cultural differences in such norms? To answer this question we used the recent Global Study of Everyday Norms (GSEN), which collected ratings of 150 scenarios in 90 societies, as the human benchmark. We prompted GPT-5 to estimate each society's average rating for every scenario, and later repeated the benchmark in three other LLMs: GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro. Compared to GSEN estimates, all four LLMs misrepresented cultural variation in two ways. First, they greatly underestimated its magnitude, estimating differences between societies to be, on average, less than half their measured size. Second, for many scenarios the LLMs poorly identified the pattern of variation, that is, which societies judged the behavior less acceptable and which societies judged it more acceptable. The pattern of variation was identified better for scenarios that elicit concerns about vulgarity, especially scenarios involving kissing and flirting. We also found that norms in more developed societies tended to be estimated somewhat more accurately, and that prompting in local survey languages rather than English produced only a modest improvement in accuracy. Local-language prompting also reduced, but did not remove, the underestimation of between-society differences. Cultural differences in everyday norms are only weakly and unevenly represented by LLMs.
发表机构
- Institute for Futures Studies(未来研究所)
- Mälardalen University(马尔默大学)
- Uppsala University(乌普萨拉大学)
- Linköping University(林雪平大学)
机构由 AI 辅助整理,请以论文原文为准。