arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.07507cs.CY

模型拒绝之处,国家恐惧之所在:威权信息控制如何在语言模型护栏中再现

What a Model Refuses, a State Fears: How Authoritarian Information Control Reproduces in Language-Model Guardrails

Menglin Liu, Yao Yu, Tong Wu, Ge Shi

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过比较十个模型和三种语言,揭示中国语言模型护栏反映威权审查逻辑,针对集体行动提示的拒绝具有选择性且易被对抗性改写绕过,表明直接衡量拒绝会高估模型的实际控制程度。

中文摘要 AI 辅助

随着大型语言模型成为政治信息的入口,它们拒绝讨论的内容成为了一种新的信息控制工具。我们认为,模型的护栏所编码的并非普遍的危害概念,而是管理其开发者所在国家的政治威胁模型,并且我们从威权政权审查方式的比较研究中推导出这种控制的预期结构。在十个模型和三种语言中,中国护栏带有其标志性特征:它们回应开发者所在政权,当提示词提到中国时,拒绝相同集体行动提示的频率远高于提到外国时;在政治领域内,它们针对的是协调能力而非异议,甚至拒绝帮助组织亲政府动员;其严格性具有渗透性,在对抗性改写下崩溃,因此最能抵抗攻击的模型是西方前沿系统,而非最严格的拒绝者。机器审查因此再现了前代信息控制基于摩擦的逻辑,但由于缺乏审查员逐案判断的能力,它比其所模仿的官僚机构更为生硬——因此直接衡量拒绝行为的审计高估了模型实际受控制的程度。

英文摘要

As large language models increasingly mediate access to political information, their refusal behavior creates a new site of information control. We ask whether political distinctions identified in theories of authoritarian information control remain visible in model guardrails, and how their enforcement changes when implemented through general-purpose AI. Across ten models, we combine cross-model comparisons with within-model experiments that manipulate political referent, collective-action potential, and political valence. Aggregate differences in political refusal across developer origins are sensitive to model composition, but the within-model experiments reveal a more specific structure. Chinese-developed models are substantially more likely to refuse otherwise identical political requests when they concern China rather than foreign or fictional settings, and refusal responds more strongly to collective-action potential than to political valence, including for peaceful pro-government mobilization. Yet these political distinctions are enforced imperfectly: refusal extends to some low-coordination political criticism even though models can distinguish criticism from coordination; the form of non-compliance varies across languages; and high native refusal does not imply adversarial robustness. These findings show that general-purpose AI can preserve recognizable political boundaries while transforming the breadth, form, and robustness with which those boundaries are enforced. More broadly, they demonstrate why refusal rates alone provide an incomplete measure of political control in language models.

↑