arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

群体保真度:评估大语言模型中的群体代表性

Population Fidelity: Evaluating Population Representativeness in LLMs

Neemias B. da Silva, Martin Lukk, Ali Sutani, Abhishek Moturu, Harris Yang, Daniel Silver, Matt Ratto, Thiago H. Silva

arXiv 2609.36253首次发表:更新:

发表机构

University of Toronto; Federal University of Technology – Paraná(多伦多大学; 巴拉那联邦理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出群体保真度评估框架,从群体准确性、变异量和变异结构三维度衡量LLM回答的群体代表性,发现文化微调提升中心对齐却未改善群体内差异表征。

AI 中文摘要

大语言模型(LLMs)在模拟人类态度和偏好方面展现出相当大的潜力。先前的研究发现,LLM生成的回答可以压缩群体内态度的分布范围,并以随模型和主题变化的方式错误表征特定子群体。我们提出了群体保真度(Population Fidelity),一个评估框架,用于区分一组LLM生成的回答表征一个群体所需的关键条件。该框架包含三个维度:群体层面的准确性、群体间变异量以及该变异的结构。我们以两种方式展示了该框架的实用性。首先,我们复现了一项关于LLM调查回答中“机器偏见”的先前研究,并将该框架应用于其模型及更新的模型,表明较差的表征不仅反映了群体间变异不足,还反映了变异被分配到错误的群体。其次,我们评估了一种旨在提高模型群体代表性的方法:文化微调(cultural fine-tuning)。我们发现文化微调可以提高与调查中心的对齐度,但并未改善群体内差异的表征,这一区别是聚合一致性指标无法捕捉的。我们认为,表征一个群体要求模型同时再现人类态度变异的多个特征。我们的框架组织了这些特征,并提供了可复用的代码、数据和训练模型,用于评估跨实质性领域的群体保真度,并评估所提出的对齐方法。

英文摘要

Large language models (LLMs) show considerable potential in simulating human attitudes and preferences. Prior work finds that LLM-generated responses can compress the range of attitudes found within populations and misrepresent particular subgroups in ways that vary across models and topics. We introduce Population Fidelity, an evaluation framework that distinguishes key conditions required for a set of LLM-generated responses to represent a population. It incorporates three dimensions: group-level accuracy, the amount of between-group variation, and the structure of that variation. We demonstrate the framework's utility in two ways. First, we reproduce a prior study of "machine bias" in LLM survey responses and apply the framework to its models and more recent ones, showing that poor representation reflects not only insufficient between-group variation but also variation assigned to the wrong groups. Second, we evaluate one proposed approach to improving models' population representativeness: cultural fine-tuning. We find that cultural fine-tuning can improve alignment with the survey center without improving the representation of within-population differences, a distinction that measures of aggregate agreement do not capture. We argue that representing a population requires models to reproduce several features of human attitudinal variation simultaneously. Our framework organizes these features and provides reusable code, data, and trained models for evaluating population fidelity across substantive domains and assessing proposed alignment methods.

Comments37 pages, 16 figures, 14 tables. Code and data: https://github.com/CriticalMaking/LLM-population-fidelity

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑