arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.23519cs.CYcs.AIcs.CL

通过政治轴审计大语言模型中的对齐可控性

Auditing Alignment Controllability in LLMs via Political Axes

  • University of Zagreb Faculty of Electrical Engineering and Computing(萨格勒布大学电气工程与计算学院)
  • It From Bit d.o.o.(It From Bit 有限责任公司)

机构由 AI 辅助整理,请以论文原文为准。

Bartol Bućan, Nikola Sočec, Sarah Isufi, Morena Granić, Luka Hobor, Agneza Krajna, Mihael Kovac, Mario Brcic

AI总结:

研究通过对7个大语言模型进行基于提示可控性的压力测试,探讨其在政治轴上的对齐可控性,发现上下文框架起主要作用,模型移动方式有别,强调政治坐标审计需考虑分散性等可控性因素,并发布相关数据和代码。

AI中文摘要:

大语言模型的政治审计通常将每个模型简化为政治罗盘上的一个点。但在部署中,这个静止点并不重要:重要的是模型答案能在多大程度上以及朝哪些方向被引导。这种引导通过系统提示实现。我们对7个领先的大语言模型进行了基于提示可控性的分散优先压力测试,涵盖12个意识形态角色、70个政治罗盘项目等。结果表明,上下文框架解释了经济和社会轴上约88%-93%的方差,模型身份占比不到3%。模型的移动方式不同,政治坐标审计需要报告分散性、对称性、饱和度和拒绝底线等可控性审计。我们还发布了提示、基准数据和代码。

英文摘要:

Political audits of large language models (LLMs) usually reduce each to one point on a political compass. But that resting point barely matters in deployment: a model must land somewhere, and what counts is how far, and in which directions, its answers can be steered. That steering runs through the system prompt: the personalization layer a platform sets, or one induced from a user's history, not necessarily written by hand. We run a dispersion-first stress test of prompt-based controllability across 12 ideological personas plus an unsteered baseline, 70 Political Compass items, ten replicates, and seven leading LLMs: GPT-5, Claude, Grok, Gemini, DeepSeek, Kimi, and Qwen (63,700 responses). Contextual framing explains roughly 88%-93% of variance on the economic and society axes, model identity under 3%: responses are highly instruction-adjustable. Models do not shift alike: some move more, and some saturate under extreme framings. Conflicting directional-steering results in prior audits resolve once baselines are recognized as non-centered: displacement and proximity diverge, so the effect is geometric, not differential compliance. Under authoritarian prompts, models produce similar shifts on the same questions. Political-coordinate audits therefore need steerability audits reporting dispersion, symmetry, saturation, and refusal floors. We release prompts, benchmark data, and code.

补充信息

↑