AI 中文总结
研究发现微调语言模型会引发意识形态泛化,提出衡量广度和放大率的方法,指出少样本提示表明泛化方向,微调会使模型走向极端,该效果能复现且对模型准确率影响小。
AI 中文摘要
在小型精选数据集上微调语言模型是使其适应特定政策或领域的标准做法。我们发现,在狭窄、事实可辩护且通过审核的数据上进行微调,会在保留一般能力的同时,导致跨不相关领域的广泛意识形态转变。在左右倾经济学问答上训练GPT-4.1,会在刑事司法、环境和文化品味等主题上产生相应的意识形态转变。同样的效果也出现在工作场所人力资源政策和实际金融查询等合理部署的数据集上,以及在科学与伪科学轴上,食品安全微调会增加对表达错误健康信念的用户的谄媚认同。我们将此现象称为意识形态泛化,并提出一种方法来衡量两个属性:广度,即转变在训练中未出现的主题上的延伸程度;放大率,即微调相对于对相同示例的少样本提示,使转变加剧的程度。我们表明,少样本提示表明泛化方向,但微调会将模型推向更远的极端,包括产生如支持种族与智商关联和政治暴力等分布外输出。该效果在Gemma-3上也能复现,在无评判评估和外部基准下成立,在与通用数据混合时依然存在,且使GSM8K准确率与基线相差在±1个百分点以内。
英文摘要
Finetuning language models on small, curated datasets is standard practice for adapting them to specific policies or domains. We show that finetuning on narrow, factually-defensible, moderation-passing data can cause broad ideological shifts across unrelated domains, while preserving general capabilities. Training GPT-4.1 on right- or left-leaning economics Q&A yields matched ideological shifts on topics such as criminal justice, the environment, and cultural taste. The same effect appears with plausibly-deployed datasets such as workplace HR policy and practical finance queries, as well as on a science-pseudoscience axis where food-safety finetuning increases sycophantic agreement with users expressing false health beliefs. We call this phenomenon ideological generalisation and propose a methodology to measure two properties: breadth, how far the shift reaches across topics absent from training, and amplification, how much finetuning intensifies the shift relative to few-shot prompting on the same examples. We show that few-shot prompting indicates the direction of generalisation but finetuning pushes the model to further extremes, including to far out-of-distribution outputs such as endorsements of race-IQ connections and political violence. The effect replicates on Gemma-3, holds under judge-free evaluations and external benchmarks, survives mixing with generic data, and leaves GSM8K accuracy within $\pm 1$pp of the baseline.