通过政治轴审计大语言模型中的对齐可控性
Auditing Alignment Controllability in LLMs via Political Axes
- University of Zagreb Faculty of Electrical Engineering and Computing(萨格勒布大学电气工程与计算学院)
- It From Bit d.o.o.(It From Bit 有限责任公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究通过对7个大语言模型进行基于提示可控性的压力测试,探讨其在政治轴上的对齐可控性,发现上下文框架起主要作用,模型移动方式有别,强调政治坐标审计需考虑分散性等可控性因素,并发布相关数据和代码。
AI中文摘要:
大语言模型的政治审计通常将每个模型简化为政治罗盘上的一个点。但在部署中,这个静止点并不重要:重要的是模型答案能在多大程度上以及朝哪些方向被引导。这种引导通过系统提示实现。我们对7个领先的大语言模型进行了基于提示可控性的分散优先压力测试,涵盖12个意识形态角色、70个政治罗盘项目等。结果表明,上下文框架解释了经济和社会轴上约88%-93%的方差,模型身份占比不到3%。模型的移动方式不同,政治坐标审计需要报告分散性、对称性、饱和度和拒绝底线等可控性审计。我们还发布了提示、基准数据和代码。
英文摘要:
Political audits of large language models (LLMs) usually reduce each to one point on a political compass. But that resting point barely matters in deployment: a model must land somewhere, and what counts is how far, and in which directions, its answers can be steered. That steering runs through the system prompt: the personalization layer a platform sets, or one induced from a user's history, not necessarily written by hand. We run a dispersion-first stress test of prompt-based controllability across 12 ideological personas plus an unsteered baseline, 70 Political Compass items, ten replicates, and seven leading LLMs: GPT-5, Claude, Grok, Gemini, DeepSeek, Kimi, and Qwen (63,700 responses). Contextual framing explains roughly 88%-93% of variance on the economic and society axes, model identity under 3%: responses are highly instruction-adjustable. Models do not shift alike: some move more, and some saturate under extreme framings. Conflicting directional-steering results in prior audits resolve once baselines are recognized as non-centered: displacement and proximity diverge, so the effect is geometric, not differential compliance. Under authoritarian prompts, models produce similar shifts on the same questions. Political-coordinate audits therefore need steerability audits reporting dispersion, symmetry, saturation, and refusal floors. We release prompts, benchmark data, and code.