扩展规模会改善基于LLM的社会模拟吗?
Will Scaling Improve Social Simulation with LLMs?
浏览论文内容
中文总结 AI 辅助
研究LLM规模扩展对社会模拟保真度的影响,发现多数任务随规模提升而改善,但存在例外,如纵向预测和低资源领域改进缓慢。
中文摘要 AI 辅助
大语言模型(LLM)社会模拟是一种有前景的研究方法,但其保真度尚不足以被广泛采用。本文研究当前语言建模中的扩展范式是否可能弥合这些差距,或者模拟保真度是否与通用能力正交,从而值得更多研究关注。我们利用扩展定律研究LLM的计算规模、通用能力基准与三个代表性子领域(观点建模、行为模拟和纵向预测)中社会模拟保真度之间的关系。令人惊讶的是,我们使用一套85个具有Qwen3架构的Transformer LLM,在固定计算预算($10^{18}$到$10^{20}$ FLOPs)下在DCLM网络文本语料库上进行预训练,在所有三个设置中发现了强大的计算扩展性。然后,我们评估了35个更大、能力更强的开放权重模型(参数高达70B),从而能够从损失预测下游准确性。这表明,大多数行为模拟和观点模拟任务将随着规模扩大而迅速改进,特别是当涉及在英文网络语料库中代表性良好的人群时。纵向预测和代表性不足的观点扩展较慢,尤其是当它们与通用知识和推理基准(如MMLU)相关性较低时。在行为模拟中,扩展未能改善模型与人类认知偏差(如风险厌恶)以及人类启发式(如从相关任务中学习相关奖励)的校准。在这些任务上,即使是微调模型,从0.5B到8B参数也未能显著提升性能。综合来看,我们得出结论:规模将在大多数设置中改善社会模拟,但存在异常值,并且在低资源领域改进将不太可靠。
英文摘要
Large Language Model (LLM) social simulations are a promising research method, but they are not yet faithful enough to be adopted widely. In this work, we investigate whether the current scaling paradigm in language modeling is likely to close these gaps, or whether simulation fidelity is orthogonal to general capabilities and therefore deserving of more research attention. We use scaling laws to study the relationship between LLMs' compute scale, general capability benchmarks, and the fidelity of social simulation in three representative sub-domains: opinion modeling, behavioral simulation, and longitudinal forecasting. Surprisingly, we discover strong compute scaling in all three settings, using a suite of 85 transformer LLMs with the Qwen3 architecture pre-trained on the DCLM web text corpus under fixed-compute budgets from $10^{18}$ to $10^{20}$ FLOPs. Then we evaluate 35 larger and more capable open-weight models up to 70B parameters, allowing us to predict downstream accuracy from loss. This reveals that the majority of behavioral and opinion simulation tasks will rapidly improve with scale, particularly when they involve populations that are well-represented in English web corpora. Longitudinal forecasting and underrepresented opinions scale more slowly, especially when they are less correlated with general knowledge and reasoning benchmarks like MMLU. In behavior simulation, scaling fails to improve model calibration with human cognitive biases like risk aversion, as well as human heuristics like learning correlated rewards from related tasks. On these tasks, even fine-tuned models fail to noticeably scale up performance from 0.5B to 8B parameters. Taken together, we conclude that scale will improve social simulations in most settings, but outliers exist, and improvements will be less reliable in low-resource domains.
发表机构
- Stanford University(斯坦福大学)
- Open Athena(开放雅典娜)
机构由 AI 辅助整理,请以论文原文为准。