arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

导航数字光谱:评估大型语言模型中的政治偏见、稳定性与下游公平性

Navigating the digital spectrum: Assessing political bias, stability, and downstream fairness in Large Language Models

Luka Debevc, Nishan Chatterjee, Antoine Doucet, Senja Pollak, Matej Martinc

arXiv 2609.08637首次发表:更新:

发表机构

Jožef Stefan Institute; Jožef Stefan International Postgraduate School; University of La Rochelle; University of Ljubljana(约热·斯泰凡研究所; 约热·斯泰凡国际研究生学院; 拉罗谢尔大学; 卢布尔雅那大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出政治罗盘测试框架,通过八维扰动空间采样评估八个大型语言模型的政治坐标,发现指令和语言显著影响结果,且角色提示对下游任务影响有限。

AI 中文摘要

大型语言模型日益被部署为信息中介,然而衡量其政治行为仍然脆弱,因为问卷结果混合了模型倾向与测量伪影及响应诱发偏差。我们引入了一个稳健的政治罗盘测试评估框架,该框架在涵盖语言、措辞、指令、答案格式、选项顺序和角色设定的八维扰动空间中采样了300种配置。我们评估了Gemma 3和Qwen 3的八个模型,涵盖14种语言和三个量化级别,获得了具有量化不确定性的设计平均政治坐标。大多数模型平均倾向于自由意志主义-左翼,但指令措辞、语言和答案格式显著影响恢复的坐标。跨语言差异主要反映坐标漂移而非不同的文化推理。对测试进行逆向工程还揭示了轴权重不平衡以及退化响应向中心坍缩的问题,因此最小模型的近原点估计可能反映弱信号而非中间主义。自由文本推理和先聊天后分类的诱发方式改变了恢复的坐标,且较大模型显示出更清晰的角色分离,其中威权主义-左翼角色在使大多数模型朝预期社会方向移动方面存在特定失败。在下游任务中,相对于模型大小和目标群体,角色效应对仇恨言论检测的影响适中,而基础提示和中间派提示在主题级情感上给出最高一致性。因此,政治角色提示具有可测量但任务和数据集特定的下游效应。

英文摘要

Large Language Models are increasingly deployed as information intermediaries, yet measuring their political behavior remains fragile because questionnaire results mix model dispositions with measurement artifacts and response-elicitation biases. We introduce a robust Political Compass Test evaluation framework that samples 300 configurations across an eight-dimensional perturbation space varying language, framing, instructions, answer format, option order, and persona wording. We evaluate eight Gemma 3 and Qwen 3 models across 14 languages and three quantization levels, obtaining design-averaged political coordinates with quantified uncertainty. Most models lean Libertarian-Left on average, but instruction phrasing, language, and answer format significantly affect recovered coordinates. Cross-lingual differences primarily reflect coordinate drift rather than distinct cultural reasoning. Reverse-engineering the test also exposes axis-weighting imbalances and the collapse of degenerate responses toward the center, so near-origin estimates for the smallest models can reflect weak signal rather than centrism. Free-text reasoning and chat-then-classify elicitation alter recovered coordinates, and larger models show clearer persona separation, with a specific failure of the Authoritarian-Left persona to move most models in the intended social direction. In downstream tasks, persona effects are modest relative to model size and target group for hate-speech detection, while base and centrist prompts give the highest agreement for topic-level sentiment. Political role prompting therefore has measurable but task- and dataset-specific downstream effects.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑