arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.02990cs.SIcs.AI

面向参与式民主的偏好推断的集体中心评估

Toward Collective-Centric Evaluation of Preference Inference for Participatory Democracy

  • Sorbonne Université(索邦大学)
  • CNRS(法国国家科学研究中心)
  • Sciences Po(巴黎政治学院)

机构由 AI 辅助整理,请以论文原文为准。

Pierre-Antoine Lequeu, Salim Hafid, Paul Lerner, Nazanin Shafiabadi, Laurène Cave, David Mas, Jean-Philippe Cointet, Benjamin Piwowarski, François Yvon

AI总结:

该研究针对参与式民主的偏好推断任务,构建集体中心评估框架并发布含9万参与者的多语言数据集,发现仅预测准确率不足以评估PI,需关注其对集体偏好结构的保留能力。

AI中文摘要:

为了扩大集体决策的规模,Polis和Remesh等参与式民主平台支持数千名参与者进行在线讨论。然而,在如此大规模下,参与者无法审查他人提交的每一条意见,产生的投票数据极为稀疏,会错误呈现共识、冲突和少数群体支持的模式。因此,这些平台越来越依赖偏好推断(Preference Inference,PI)模型来预测缺失的投票。但这种自动化并非中立:推断出的偏好可能人为放大、抑制或重新排序现有的支持模式,最终改变对讨论结果的解读。更广泛地说,我们缺乏对现有PI方法如何影响集体偏好格局的系统理解。为解决这一缺口,我们在该场景下对几种现有PI方法进行了基准测试。我们超越了以个体预测准确性为核心的传统用户中心评估,引入了一种集体中心评估框架,用于衡量推断投票是否保留更广泛偏好格局的显著属性。我们还贡献了同类中最大的多语言数据集:四次咨询涵盖超过9万名参与者、100万次投票和22种语言。实验表明,预测准确率相当的模型在保留集体结构的程度上存在显著差异。这些结果证明,仅靠准确率不足以评估民主场景下的PI。这项工作通过为PI任务提供一种新颖、全面且集体中心的评估基准,旨在支持AI系统的开发,这些系统在扩大讨论规模的同时不损害其民主结果的完整性。

英文摘要:

To scale up collective decision-making, participatory democracy platforms such as Polis and Remesh enable online deliberation among thousands of participants. However, at this scale, participants cannot review every opinion submitted by others, producing highly sparse voting data that misrepresent patterns of consensus, conflict, and minority support. Platforms therefore increasingly rely on Preference Inference (PI) models to predict missing votes. Yet this automation is not neutral: inferred preferences can artificially amplify, suppress, or reorder existing patterns of support, ultimately reshaping how the outcomes of a deliberation are interpreted. More generally, we lack a systematic understanding of how existing PI methods affect the collective preference landscape. To address this gap, we benchmark several existing PI approaches in this context. Moving beyond conventional user-centric evaluations centered on the accuracy of individual predictions, we introduce a collective-centric evaluation framework that measures whether inferred votes preserve salient properties of the broader preference landscape. We further contribute the largest multilingual dataset of its kind: four consultations spanning over 90k participants, 1M votes, and 22 languages. Our experiments show that models with comparable predictive accuracy can differ substantially in the degree to which they preserve the collective structure. These results demonstrate that accuracy alone is insufficient for evaluating PI in democratic settings. By contributing a novel comprehensive and collective-centric evaluation benchmark for the task of PI, this work aims to support the development of AI systems that scale deliberation without compromising the integrity of its democratic outcomes.

补充信息

↑