arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

政治分化可通过用户反馈驱动AI模型彼此分离

Political Sorting Can Drive AI Models Apart Through User Feedback

Petter Törnberg, Michael Heseltine, Nicolò Pagan, Christopher Bail, Michelle Schimmel, Christopher Barrie

arXiv 2609.28486首次发表:更新:

发表机构

Institute for Logic, Language and Computation, University of Amsterdam; Department of Sociology, University of Oxford; Department of Informatics, University of Zurich; Department of Sociology, Duke University; Department of Sociology, New York University(阿姆斯特丹大学逻辑、语言与计算研究所; 牛津大学社会学系; 苏黎世大学信息学系; 杜克大学社会学系; 纽约大学社会学系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究揭示政治分化可通过用户反馈驱动AI模型碎片化,提出“离心对齐螺旋”机制,并通过实验与仿真证明其能导致模型间持久政治差异。

AI 中文摘要

大型语言模型正迅速成为政治信息的重要来源。这引发了一个根本性问题:AI系统是会支持政治知识的共享基础,还是会导致不同政治群体依赖日益不同的模型?如果三个条件成立,政治分化就能驱动模型碎片化:政治立场不同的用户选择不同的模型,从用户反馈中学习会使这些模型在政治上彼此分离,且由此产生的差异会塑造后续的模型选择。我们将这一自我强化的过程称为“离心对齐螺旋”。我们分三步研究其组成部分。首先,我们利用一项人类实验表明,政治身份能预测模型选择。其次,我们在反映 predominantly 民主党或共和党偏好的合成反馈上对语言模型进行微调。在每个模型族的五次独立运行中,配对模型在12-41%的未见调查问题上出现分歧,且存在巨大的党派差距,并且在每次运行中,差异都朝着预期的政治方向移动;对于某些模型,分化甚至扩展到训练中未包含的议题领域。相反,跨群体汇总反馈则抑制了分歧。第三,一个经验锚定的基于智能体的模型展示了当政治分化和模型适应共同作用时会发生什么:模型吸引政治立场不同的受众,从他们那里学习,并进一步分化。因此,用户反馈可以将AI用户中的政治分化转化为他们依赖获取政治信息的模型之间的持久差异。

英文摘要

Large language models are rapidly becoming an important source of political information. This raises a fundamental question: will AI systems support a shared basis for political knowledge, or lead different political groups to rely on increasingly different models? Political sorting can drive model fragmentation if three conditions hold: politically different users select into different models, learning from user feedback pushes those models apart politically, and the resulting differences shape subsequent model choices. We call this self-reinforcing process the Centrifugal Alignment Spiral. We study its components in three steps. First, we draw on a human experiment showing that political identity predicts model choice. Second, we fine-tune language models on synthetic feedback reflecting predominantly Democratic or Republican preferences. Across five independent runs per model family, paired models diverged on 12-41% of unseen survey questions with large partisan gaps, and in every run the differences moved in the expected political direction; for some models, differentiation extended even to issue areas excluded from training. Pooling feedback across groups instead suppressed divergence. Third, an empirically anchored agent-based model shows what follows when political sorting and model adaptation operate together: models attract politically distinct audiences, learn from them, and diverge further. User feedback can therefore turn political sorting among AI users into durable differences between the models on which they rely for political information.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑