面向数字与AI服务使用的韩国合成角色面板的分布有效性与校准:针对韩国媒体面板调查的二次数据验证
Distributional Validity of a Korean Synthetic Persona Panel: Evidence From the Korea Media Panel Survey
浏览论文内容
中文总结 AI 辅助
本研究验证韩国合成角色面板NVIDIA Nemotron-Personas-Korea的分布有效性与校准,发现其无法替代真实调查数据,仅在真实数据极稀缺或用于诊断场景时具备价值。
中文摘要 AI 辅助
基于大型语言模型(LLMs)的合成角色日益被提议作为人类调查受访者的替代品,但英语语境之外的系统验证仍然稀缺。本项二次数据研究评估了韩国合成角色面板NVIDIA Nemotron-Personas-Korea的表现,该面板分别基于Gemini 3.5 Flash(主模型)和EXAONE(对比模型)构建,在多大程度上复现KISDI韩国媒体面板调查中的数字与AI服务使用分布。每个模型构建了按性别和年龄分层的约8000个角色的面板,回答调查自身的8项服务使用指标以及8项创新性与接受度构念,并与加权调查估计值进行比较。总体平均绝对误差(MAE,对应研究问题1)为15-19个百分点(pp),二元项均值相关性为0.69-0.90;5个人口统计维度的细分误差(对应研究问题2)为15-19个百分点,组间差异最高达52.4/36.2个百分点(Gemini/EXAONE)。误差呈现模型特有的特征:Gemini存在锚定不足的年龄刻板印象,EXAONE存在与默许倾向一致的水平偏差。参考年份分析表明,时间错位是生成式AI高估的主要原因,而短格式低估受框架影响。对30%真实数据的留存校准(对应研究问题3)大致将性别-年龄单元格MAE减半(从18.9降至8.6、15.9降至6.7个百分点),但从相同真实子样本直接估计的准确性要高得多(3.6个百分点),且该校准无法跨时间迁移。校准后的面板仅在真实数据极稀缺(约100个响应)的情况下保留优势,且对其中一个模型而言,适用于未观测细分群体;角色叙事条件化优于仅人口统计条件化,但均未超过简单真实数据基线。因此,合成面板并非调查替代品,其价值在于诊断,实际使用仅限于缺乏真实数据的场景。
英文摘要
Large language model (LLM) personas are proposed as survey respondents, yet validation outside English-speaking contexts is scarce. We evaluate how well a Korean synthetic persona panel used to condition Gemini 3.5 Flash and EXAONE reproduces digital and artificial intelligence (AI) service-use distributions of the Korea Media Panel Survey. About 8,000 personas per model answered eight service-use items and eight attitudinal constructs; responses were compared with weighted survey estimates. The overall mean absolute error (MAE) was 14-19 percentage points (pp), with binary item-mean correlations of 0.70-0.91 across waves. Segment error across five axes was 14-18 pp, with between-group signed-error ranges of 49.6/34.7 pp (Gemini/EXAONE; 39.5/31.2 without the non-comparable teen cells). Errors were model-specific: an age stereotype (Gemini) versus an acquiescence-consistent level bias (EXAONE). Generative-AI overestimation was consistent with temporal misalignment; short-form underestimation was framing-sensitive and persisted under randomized order (both shown for Gemini). Post-hoc holdout calibration on 30% of the real data, with the correction form selected inside the calibration set, cut cell MAE from 18.3/15.1 to 4.9/4.4 pp, yet direct estimation from that subsample was more accurate than the calibrated panel (3.6 pp), a synthetic-informed shrinkage estimator beat its real-only counterpart by at most 0.7 pp, and the correction did not transfer competitively across waves. The calibrated panel kept an advantage only below roughly 250-860 real responses (at most 2.3 pp over a real-only shrinkage estimator) or, for one model, on unobserved segments. In this setting, synthetic panels are not survey substitutes; their value is diagnostic.
发表机构
- College of AI Convergence, Seoul Cyber University(首尔网络大学人工智能融合学院)
- Sungkyunkwan University(成均馆大学)
机构由 AI 辅助整理,请以论文原文为准。