发表机构
Korea University(韩国大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究诊断LLM合成角色数据的人口统计分布与外部参考的匹配度,发现误差多源于参考选择,采用rake加权等方法调整后偏差降低,建议将合成角色数据作为调查辅助材料。
AI 中文摘要
我们诊断基于大语言模型(LLM)的合成角色数据中的人口统计分布与外部参考分布的匹配程度。针对所考察的三个变量,我们表明,观测到的误差大部分可归因于参考的选择,而非生成器。使用总变差距离(TVD),我们将Nemotron-Personas-Korea(NPK)的100万条记录的性别×年龄组×省份联合分布与韩国官方统计数据进行比较。针对使用时间即2026年4月的居民登记数据,偏差边界(由三个变量构成的任何子组占比的最大可能差异)为1.81个百分点,这与约2900名受访者的调查误差幅度相当。该值并非数据的固定属性,将参考时期和序列与此处确定的生成参考匹配后,2024年仅限韩国公民的登记人口普查将偏差边界降至0.56个百分点。在匹配度最佳的月份(2025年1月)与使用时间之间的15个月里,居民登记人口结构自身的变动幅度超过NPK最小误差的两倍。韩国调查实践中使用的两种加权方案—— rake( rake加权)和单元格事后分层,在两种情况下均以约0.2%的方差膨胀消除了大部分参考时期依赖性。针对生成参考进行rake加权后,剩余的联合差异处于且略高于完美生成器生成100万条记录时会产生的差异上限(蒙特卡洛分布的第97.6百分位)。我们建议将合成角色数据视为小规模调查设计的辅助材料,而非调查数据的替代品,并针对使用时的官方统计数据重新进行诊断和调整。
英文摘要
We examine how strongly demographic-distribution diagnostics of LLM-based synthetic persona data depend on the official statistics chosen as the reference. We compared the sex $\times$ age-group $\times$ province joint distribution of the 1,000,000 Nemotron-Personas-Korea (NPK) records with Korean official statistics using total variation distance (TVD). Against the resident-registration population for April 2026, the time of use, the upper bound on the share discrepancy of any subgroup defined by the three variables was 1.81 percentage points, the margin of error of a survey with about 2,900 respondents. This value, however, depended on the reference date and population definition of the official statistics. The closest candidate examined, the 2024 census distribution for Korean nationals, gave 0.56 percentage points; no monthly resident-registration reference was closer. The reference distributions themselves also moved: between January 2025, when NPK's distance was smallest, and April 2026, the resident-registration distribution moved by more than twice that minimum distance, as did the census Korean-national distribution between 2024 and 2025. Raking reduced the range of the distance across reference months to about one fifth. Cell poststratification was applicable because none of the 204 cells was empty, and neither scheme produced extreme weights. Demographic alignment is necessary but not sufficient for response validity, and even this basic diagnostic depends on the choice of reference distribution. Providers should therefore disclose the source, reference date, population definition, and joint-cell allocation procedure of the official statistics used for demographic attributes, and users should repeat the diagnosis and adjustment against statistics current at the time of use.
Comments27 pages, 2 figures, 6 tables. Data, materials, and code: https://doi.org/10.7910/DVN/BUFIGC