arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19689cs.AI

重新思考基于大语言模型(LLM)的社会模拟的评估与优化

Rethinking the Evaluation and Optimization of LLM-Based Social Simulation

Pei Wang, Xu Chen, Ji-Rong Wen

首次发表
浏览论文内容

中文总结 AI 辅助

针对LLM社会模拟中基于准确率的评估与硬标签训练的缺陷,提出主观性自适应软标签训练(SALT),构建SUBJSIM基准并验证其在现实场景下的可行性与优势。

中文摘要 AI 辅助

基于大语言模型(LLM)的社会模拟是对调查、行为实验等传统方法的有前景的补充。核心问题在于如何评估LLM模拟人类行为的保真度,并针对该保真度优化LLM。当前主流实践通过准确率进行评估,即检查模型是否选择人类观察到的单一响应,并训练LLM以复现该硬标签。然而,人类行为本质上具有主观性:同一人在相同情境下可能会做出合理的不同行为,因此观察到的响应只是潜在响应分布中的一次抽样,这使得基于准确率的评估不可靠,且硬标签训练具有误导性。为解决这些问题,我们首先引入主观性系数,这是一种基于熵的量,用于区分编码等客观任务与社会模拟等主观任务,并利用该系数系统分析随着主观性增长,基于准确率的评估和硬标签训练的失效情况。基于主观性系数,我们提出主观性自适应软标签训练(SALT):它将语义相近输入的观察输出汇集为软分布标签,聚合半径适配于每个输入的估计主观性;在接近客观的极限情况下,邻域会缩小,因此SALT自然退化为标准单标签训练。此外,由于现有数据集仅记录单一观察响应,无法支持分布评估,我们构建了SUBJSIM基准,包含19300个情境,覆盖193名标注者和100个主观问题。由于现实世界数据通常仅为每个输入提供单一观察值,我们的实验从单一观察输出训练模型,同时针对完整响应分布对其进行评估,验证了在现实场景中的可行性。在SUBJSIM上的结果证明了我们方法的优势。

英文摘要

LLM-based social simulation is a promising complement to traditional methods such as surveys and behavioral experiments. A core question is how to evaluate the fidelity of LLM-simulated human behavior and optimize LLMs toward it. Prevailing practice evaluates by accuracy, checking whether the model selects the single response observed from a human, and trains the LLM to reproduce this hard label. However, human behavior is inherently subjective: the same person in the same situation may reasonably act differently, so an observed response is only one draw from an underlying response distribution, rendering accuracy-based evaluation unreliable and hard-label training misleading. To address these problems, we first introduce the subjectivity coefficient, an entropy-based quantity distinguishing objective tasks such as coding from subjective ones such as social simulation, and use it to systematically analyze how accuracy-based evaluation and hard-label training fail as subjectivity grows. Based on the subjectivity coefficient, we propose Subjectivity-Adaptive soft-Label Training (SALT): it pools observed outputs from semantically nearby inputs into soft distributional labels, with an aggregation radius adapted to the estimated subjectivity of each input; in the near-objective limit the neighborhood shrinks, so SALT naturally falls back to standard single-label training. Moreover, since existing datasets record only single observed responses and cannot support distributional evaluation, we construct SUBJSIM, a benchmark of 19,300 contexts covering 193 annotators and 100 subjective questions. Since real-world data typically provide only a single observation per input, our experiments train models from single observed outputs while evaluating them against the full response distributions, verifying feasibility in realistic settings. Results on SUBJSIM demonstrate the advantages of our method.

发表机构

  • Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学高瓴人工智能学院)

机构由 AI 辅助整理,请以论文原文为准。

↑