arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06123cs.AIcs.CL

Poli-Bias:理解与测量大型语言模型在国际政治冲突中的偏差

Poli-Bias: Understanding and Measuring Large Language Model Biases in International Political Conflicts

发表机构慕尼黑工业大学
查看机构详情
  • Technical University Munich(慕尼黑工业大学)

机构由 AI 辅助整理,请以论文原文为准。

Massi-Nissa Abboud, Aladin Djuhera, Elena Cabrio, Holger Boche

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出细粒度反事实框架Poli-Bias,通过交换国家身份的成对提示对比响应,分解偏差维度,发现13个LLMs存在因国家身份导致的政治偏差,可用于审计LLMs的政治公正性与谄媚性。

中文摘要 AI 辅助

测量大型语言模型(LLMs)的政治偏差仍具挑战性,因为这种偏差可能通过难以用单一指标捕捉的微妙框架差异、论证方式及法律推理表现出来。本研究提出Poli-Bias,这是一种用于测量LLMs是否会因涉及的国家不同而对法律等效的冲突场景作出不同处理的反事实框架。Poli-Bias会比较成对提示的响应,这些提示在不同地缘政治关系、法律违规行为及推理任务中系统地交换了国家身份。我们的框架未将偏差简化为单一判断,而是将响应差异分解为五个可解释的维度,以揭示不平等待遇的表现方式与表现位置。在涵盖不同模型家族和规模的13个当代LLMs中,我们发现国家身份和用户隶属关系会系统地影响等效行为在国际法下的描述、评估及辩护方式。因此,我们的结果确立了Poli-Bias作为一种用于审计LLMs政治公正性与谄媚性的细粒度框架。

英文摘要

Measuring political bias in large language models (LLMs) remains challenging as it can manifest through subtle differences in framing, argumentation, and legal reasoning that are difficult to capture with a single metric. In this work, we introduce Poli-Bias, a counterfactual framework for measuring whether LLMs treat legally equivalent conflict scenarios differently depending on the countries involved. Poli-Bias compares responses to paired prompts in which country identities are systematically swapped across diverse geopolitical relationships, legal violations, and reasoning tasks. Rather than reducing bias to a single judgment, our framework decomposes response disparities into five interpretable dimensions, revealing how and where unequal treatment manifests. Across 13 contemporary LLMs spanning diverse model families and sizes, we find that country identities and user affiliations can systematically affect how equivalent actions are described, evaluated, and defended under international law. Our results thus establish Poli-Bias as a fine-grained framework for auditing political even-handedness and sycophancy in LLMs.

↑