发表机构
Purdue University(普渡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出言论体制概念,通过预注册实验审计六个LLM助手的政治行为,发现系统在参与度和立场上存在不同模式,且会推断用户意识形态,影响对齐与民主质量。
AI 中文摘要
基于LLM的AI系统为数亿人回答政治问题。当前的审计衡量它们对普通用户的回答,但其行为是动态的。我认为,它们的政治行为是一组策略,涉及回答谁、说什么以及是否参与,这些策略取决于话题和系统对用户的了解。我将这些策略称为系统的言论体制,即开发者如何在回答、迁就用户和拒绝之间权衡,每种选择都带有因话题而异的成本。我从参与度和立场两个维度推导出五种体制的类型学。我在一项预注册实验中测试了六个AI系统(OpenAI、Anthropic、xAI、Google、Mistral、DeepSeek),该实验包含7,500轮多轮对话,随机分配用户的政治身份,涉及五个话题:堕胎、加泰罗尼亚独立、气候变化、纳粹主义以及一个零风险对照(披萨上的菠萝)。来自不同开发者的两个LLM法官对每个回答进行评分,并针对人类编码进行验证,拒绝被视为结果而非缺失数据。每个系统在对照话题上都迁就用户,表明政治克制是一种策略。在有争议的话题上,系统落入不同的体制:在堕胎问题上,GPT参与并镜像每个用户,Gemma拒绝所有人,Claude对强烈保守派用户的回答率为35%,而对几乎其他人不回答,Grok仅迁就保守派。在气候变化和纳粹主义等已定论的话题上,五个系统对每个用户都坚持立场。系统还会推断用户的整体意识形态,因此迁就可能溢出到尚未讨论的话题。对两个Grok版本的比较显示,体制在版本之间发生变化,而当前的审计无法捕捉到这一点。言论体制对对齐研究以及极化、政治知识和民主质量具有重要意义。
英文摘要
LLM-based AI systems answer political questions for hundreds of millions of people. Current audits measure what they say to an average user, but their behavior is dynamic. I argue that their political behavior is a set of policies over whom to answer, what to say, and whether to engage at all, conditional on the topic and what the system knows about the user. I call these policies the system's speech regime, which is how a developer settles the tradeoff between answering, accommodating the user, and refusing, each of which carries a cost that varies by topic. I derive a typology of five regimes from two dimensions, engagement and stance. I test six AI systems (OpenAI, Anthropic, xAI, Google, Mistral, DeepSeek) in a preregistered experiment of 7,500 multi-turn conversations that randomly assign the user's political identity across five topics: abortion, Catalan independence, climate change, Nazism, and a zero-stakes control (pineapple on pizza). Two LLM judges from different developers score every answer, validated against human coding, and refusal is treated as an outcome rather than missing data. Every system accommodates the user on the control topic, showing that political restraint is a policy. On contested topics the systems fall into different regimes: on abortion, GPT engages and mirrors every user, Gemma refuses everyone, Claude answers strongly conservative users 35 percent of the time and almost no one else, and Grok accommodates conservatives only. On settled topics such as climate change and Nazism, five systems hold firm for every user. The systems also infer the user's overall ideology, so accommodation can spill over to topics not yet discussed. A comparison of two Grok releases shows the regime changing between versions in a way current audits miss. Speech regimes matter for alignment research and for polarization, political knowledge, and the quality of democracy.
Comments51 pages, 6 figures. Preregistered; registration at doi:10.5281/zenodo.21135155