发表机构
Department of Politics and International Relations, University of Oxford(政治与国际关系系,牛津大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型政策评估是否受地缘政治线索影响,通过背书实验让四个模型评估相同政策,发现模型评分受背书者身份影响,西方背书被视为可信度线索,中国和俄罗斯背书被视为其他风险线索。
AI 中文摘要
大语言模型(LLMs)越来越多地用于总结和评估与政策相关的信息,但其判断是否受到地缘政治线索的影响尚不清楚。本文通过背书实验研究该问题,让四个大语言模型评估相同的国际经济和安全政策,并随机将政策描述为由美国、欧盟、中国或俄罗斯支持。在纯数字条件下,GPT-5、Claude Sonnet和Gemini对中国和俄罗斯支持的政策评分显著低于美国或欧盟支持的相同政策;DeepSeek是主要例外。在要求模型提供简短理由的条件下,GPT-5和Claude Sonnet中西方/非西方的差距依然存在,Gemini的惩罚有所减轻,而DeepSeek中对中国和俄罗斯的惩罚大幅增加。理由表明,西方的背书常被视为可信度线索,而中国和俄罗斯的背书则被视为数据安全、主权、监视或地缘政治风险的线索。这些发现表明,即使政策内容固定,大语言模型的政策评估也可能取决于外国背书者的身份。
英文摘要
Large language models (LLMs) are increasingly used to summarize and evaluate policy-relevant information, but it remains unclear whether their judgments are implicitly shaped by geopolitical cues. I study this question with an endorsement experiment in which four LLMs evaluate the same international economic and security policies after each policy is randomly described as supported by the United States, the European Union, China, or Russia. In the numeric-only condition, GPT-5, Claude Sonnet, and Gemini rate China- and Russia-endorsed policies substantially lower than identical policies endorsed by the United States or the European Union; DeepSeek is the main exception. A second condition asks models to provide a short justification with the score. This request leaves the broad Western/non-Western gap intact for GPT-5 and Claude Sonnet, attenuates Gemini's penalties, and sharply activates China and Russia penalties in DeepSeek. The justifications indicate that Western endorsement is often treated as a credibility cue, whereas Chinese and Russian endorsement is treated as a cue for data security, sovereignty, surveillance, or geopolitical risk. These findings show that LLM policy evaluations can depend on the identity of a foreign endorser even when policy content is held fixed.