arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.25953cs.CLcs.CY

Polistemics:评估大语言模型在政治与选举中作为信息中介的表现

Polistemics: Evaluating LLMs as Information Mediators in Politics & Elections

发表机构瑞士联邦理工学院智能系统实验室 · 无合适常见中文名,可音译为“科尔达”
查看机构详情
  • ETH Agentic Systems Lab(瑞士联邦理工学院智能系统实验室)
  • CORDA(无合适常见中文名,可音译为“科尔达”)

机构由 AI 辅助整理,请以论文原文为准。

Baran Peters, Gabor Hollbeck, Robert Jakob, Kevin O'Sullivan

首次发表
浏览论文内容

中文总结 AI 辅助

研究探讨如何评估大语言模型在政治选举中作为信息中介的表现。引入基于认知谦逊的Polistemics基准,在不同信息属性环境下测试。应用于三个大语言模型发现,虽总分高但有系统性故障,受党派先验影响,可靠中介可实现但模型表现不稳定。

中文摘要 AI 辅助

随着大语言模型(LLMs)越来越多地作为公民所依赖的政治信息中介,目前仍没有标准化方法来评估它们是否负责地履行这一职责。我们引入了Polistemics,这是一个基于理论的基准,用于评估大语言模型在选举中作为政治信息中介的情况。先前的工作将此任务视为复制而非中介,未解决其认知维度以及与不完美信息的交互问题。我们将评估建立在认知谦逊之上,这是一种源自公民认知能动性的规范标准,并在不同信息属性(如清晰度、噪声和一致性)的受控环境中进行测试。将该基准应用于针对2025年德国和荷兰选举的三个先进大语言模型,我们发现高总分掩盖了系统性故障。模型在明确证据下能可靠地进行中介,但在信息缺失、模糊或矛盾时会失效,同时还会使政治语言的强度趋于平缓。这些故障可能由党派先验驱动,受党派标签和输出语言影响。可靠的中介似乎是可以实现的,但没有模型能始终如一地做到。

英文摘要

As LLMs increasingly shape the political information citizens rely on, no standard exists to assess whether they do so responsibly. We introduce Polistemics, a theory-grounded diagnostic benchmark for evaluating LLMs as mediators of political information in elections. Prior work has treated this task as reproduction rather than mediation, leaving its epistemic dimensions and interaction with imperfect information unaddressed. We ground the evaluation in Epistemic Modesty, a normative standard derived from citizens' epistemic agency, and test it across controlled settings that vary the clarity, noise, and consistency of the available evidence. Applying the benchmark to three state-of-the-art LLMs across the 2025 German and Dutch elections, we find that high aggregate scores mask systematic failures. Models mediate reliably under clear evidence but break down when it is absent, vague, or contradictory, while flattening the intensity of political language throughout. These failures point to party priors, shifting with party labels and output language. Reliable mediation appears achievable, but no model delivers it consistently.

补充信息

↑