arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型文化对齐中的测量有效性

Measurement Validity in LLM Cultural Alignment

An Duy Nguyen, Muhammad Aurangzeb Ahmad

arXiv 2608.29266首次发表:更新:

发表机构

University of Washington; University of Washington Bothell(华盛顿大学; 华盛顿大学 Bothell 分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对多个LLM区分调查回复、采样噪声与问题框架,发现多数模型的表观文化位置无法与噪声区分,提示需先验证测量可靠性才能将LLM调查回复用于文化归因。

AI 中文摘要

研究人员越来越多地将大语言模型(LLM)的调查回复作为人类文化价值观的替代指标,包括将模型输出投射到英格尔哈特-韦尔策尔文化地图(Inglehart-Welzel Cultural Map)等工具上,并得出模型与哪些文化相似的结论。虽然模型对带有价值倾向问题的回答可被解释为文化信号,但它也带有采样噪声,且对问题框架相当敏感。在本文中,我们针对多个LLM区分调查回复、采样噪声和问题框架,将这些模型的回复方差分解为随机种子、提示改写产生的变异。我们采用噪声信号比(NSR)来测试模型的表观文化位置是否可与噪声区分开。当将其应用于来自四个地理来源的十几个模型(针对88个综合价值观调查国家进行校准)时,结果通常为不可区分:在117个有效模型-问题对中,有49个(42%)的NSR超过1.0,最坏情况下达到5.56;还有两个模型直接拒绝回答足够数量的调查问题。我们的结果证实了先前的发现,即LLM会向西方、英语文化位置聚类。然而,本研究中不成立的是,目前任何人能精确解释特定模型坐标的程度:仅提示语气就可使模型偏移2.4个地图单位,这与英格尔哈特-韦尔策尔文化地图中实际国家之间的距离相当。这些发现表明,在将LLM调查回复的文化归因解释为文化表征的证据之前,需要先建立基础测量的可靠性。

英文摘要

Researchers increasingly treat LLM survey responses as a proxy for human cultural values. This includes projecting model outputs onto instruments like the Inglehart-Welzel Cultural Map and drawing conclusions about which cultures a model resembles. While a model's answer to a value-laden questions may be interpreted as a cultural signal, it also carries sampling noise and, can be quite sensitive to question framing. In this paper, we separate survey responses, sampling noise and question framing for multiple LLMs. We decompose response variance from these models into variation across random seeds, prompt rewordings. We employ noise-to-signal ratio (NSR) to test whether a model's apparent cultural position is distinguishable from noise. When applied across a dozen models from four geographic origins, calibrated against 88 Integrated Values Survey countries, the answer is often no. NSR exceeds 1.0 on 49 of 117 valid model-question pairs (42%), reaching 5.56 in the worst case. Two models even refuse to answer sufficient number of survey questions outright. Our results corroborate previous findings that LLMs cluster toward Western, English-speaking cultural positions. However, what does not hold up in this study is the precision with which anyone can currently interpret a specific model's coordinates: prompt tone alone can shift a model by 2.4 map units, comparable to the distance between actual countries in the Inglehart-Welzel Cultural Map. These findings suggest that cultural attribution from LLM survey responses requires establishing the reliability of the underlying measurements before interpreting model coordinates as evidence of cultural representation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑