关注差距:用于人类模拟的思维混合模型
Mind the Gaps: Mixture-of-Minds for Human Simulation
浏览论文内容
中文总结 AI 辅助
本文提出Anacreon受众模拟模型,基于Gemma 4 12B构建思维混合模型,通过聚类个体、情感链等技术,在个体层面序数对齐达0.775,缩小了群体与个体模拟的差距。
中文摘要 AI 辅助
预测群体如何回答新问题是一个长期目标。统计方法在群体层面取得成功,但在个体层面表现不佳。大型语言模型模拟器继承了这一差距:它们恢复了群体的中心趋势,却抹平了异质性,还带有社会偏见和提示脆弱性,扭曲了个体预测。本文提出Anacreon,一个针对狭窄、明确领域内个体层面的受众模拟模型。Anacreon学习将个体区分开的作者嵌入,围绕种子人物聚类真实定性语料库,并基于Gemma 4 12B基础模型为每个聚类训练专用适配器,即思维混合模型。它从公开文本中获取人口统计数据、心理特征和调查响应,并用情感链扩充每条记录。它通过打乱响应选项降低提示脆弱性,通过平衡训练分布降低正向偏见。在大型外部调查中,Anacreon达到了该领域已达成共识的个体层面准确率指标序数对齐0.775,残留偏见较小。该研究是从忠实模拟的个体中提取群体见解的一步。
英文摘要
Predicting how a population will answer a new question is a long-standing goal. Statistical methods succeed at the level of the mass but falter at the level of the individual. Large language model simulators inherit this gap. They recover a population's central tendencies while flattening its heterogeneity, and they carry social biases and prompt brittleness that distort individual predictions. This paper introduces Anacreon, an audience simulation model that targets the individual level within a narrow, well-specified domain. Anacreon learns an authorship embedding that separates individuals, clusters a real qualitative corpus around seed people, and trains a dedicated adapter for each cluster, a mixture of minds, on a Gemma~4 12B base. It harvests demographics, psychological traits, and survey responses from public text, and augments each record with a chain-of-emotion. It reduces prompt brittleness by shuffling response options and reduces positive bias by balancing the training distribution. On a large, externally sourced survey, Anacreon reaches a state-of-the-art ordinal alignment of 0.775, the individual-level accuracy measure on which the field has converged, with a small residual bias. The work is a step toward drawing aggregate insight from faithfully simulated individuals.
发表机构
- Semilattice(半格)
机构由 AI 辅助整理,请以论文原文为准。