从联合国投票数据估计大语言模型的地缘政治偏好
Estimating the Geopolitical Preferences of Large Language Models from United Nations Voting Data
浏览论文内容
中文总结 AI 辅助
研究如何从联合国投票数据估计大语言模型的地缘政治偏好,采用动态序数理想点方法,将模型视为对相关决议全文的回应者,得出不同模型支持率及与五常关系等结果,发现模型地缘政治立场与开发者母国可能不同。
中文摘要 AI 辅助
研究人员应如何衡量大语言模型(LLMs)表达的地缘政治偏好?现有审计通常依赖调查和简单测试,但国际关系研究早就认识到衡量地缘政治偏好很困难,并已开发出从观察到的选择中恢复偏好的方法。本文应用国际关系中的动态序数理想点方法,将大语言模型视为对1946年至2025年联合国大会常会审议的5555项有分歧、有记录、已通过决议全文的回应者。支持率从DeepSeek的37.8%到GPT - 5的97.3%不等。令人惊讶的是,在21世纪,GPT - 5、Claude Sonnet和Gemini在五常中最接近俄罗斯;DeepSeek最接近法国;且这四个都离美国最远。在2104项美国反对但中国和俄罗斯/苏联支持的决议中,GPT - 5支持96.1%,Gemini支持83.4%,Claude Sonnet支持65.2%,DeepSeek支持36. | 1%。研究结果表明,模型表达的地缘政治立场可能与其开发者母国的立场明显不同,尤其是在国际政治中,国家行动可能与模型训练文本中普遍存在的既定原则不同。
英文摘要
How should researchers measure the geopolitical preferences expressed by large language models (LLMs)? Existing audits commonly rely on surveys and simple tests, but international-relations research has long recognized that measuring geopolitical preferences is difficult and has developed methods for recovering them from observed choices. This paper applies a dynamic ordinal ideal-point approach from international relations, treating LLMs as respondents to the full texts of 5,555 divisive, recorded, adopted resolutions considered in regular sessions of the UN General Assembly from 1946 through 2025. Support ranges from 37.8% for DeepSeek to 97.3% for GPT-5. Surprisingly, in the twenty-first century, GPT-5, Claude Sonnet, and Gemini are closest among the permanent five to Russia; DeepSeek is closest to France; and all four are farthest from the United States. Among 2,104 resolutions opposed by the United States but supported by China and Russia/USSR, GPT-5 supported 96.1%, Gemini 83.4%, Claude Sonnet 65.2%, and DeepSeek 36.1%. The findings show that a model's expressed geopolitical position can differ markedly from that of its developer's home country, especially in international politics, where state actions can diverge from the stated principles prevalent in the texts on which models are trained.