发表机构
Korea Advanced Institute of Science and Technology(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出角色条件化信息性(PCI)指标,通过将角色建模为相似度图并利用Local Moran's I量化一致性,筛选合成受访者,在PVQ-RR数据集上验证其能提升潜在价值观结构的恢复效果。
AI 中文摘要
角色条件化的大语言模型(LLMs)正越来越多地被用于模拟不同领域的调查回复。然而,明显的回复差异可能反映的是未条件化的模型先验或token采样噪声,而非系统性的角色条件化。我们提出,当语义相似的角色表现出一致的回复变化时,角色条件化的差异具有信息性。为将该原则付诸实践,我们引入了角色条件化信息性(PCI),这是一种无监督诊断指标,用于衡量语义相似的角色是否相对于项目级样本基线向一致方向偏离。通过将角色建模为相似度图,PCI利用Local Moran's I量化局部空间一致性并提取紧凑的角色子集,且无需使用构念标签。为在不依赖外部人类基准的情况下评估PCI,我们使用包含57个项目的肖像价值观问卷修订版(PVQ-RR)测试其恢复已确立的潜在价值观结构的能力。验证性因子分析(CFA)显示,PCI选出的10%子集相较于回复稳定性和随机选择,显著提升了整体构念恢复效果。这些发现支持PCI作为调查流程中筛选合成受访者的原则性内部诊断工具。
英文摘要
Persona-conditioned large language models (LLMs) are increasingly used to simulate survey responses across diverse domains. However, apparent response variation can reflect unconditioned model priors or token sampling noise rather than systematic persona conditioning. We argue that persona-conditioned variation is informative when semantically similar personas exhibit concordant response shifts. To operationalize this principle, we introduce Persona-Conditioned Informativeness (PCI), an unsupervised diagnostic metric that measures whether semantically similar personas deviate in concordant directions relative to item-level sample baselines. By modeling personas as a similarity graph, PCI uses Local Moran's I to quantify local spatial coherence and extract compact persona subsets without using construct labels. To evaluate PCI without external human benchmarks, we test its ability to recover established latent value structure using the 57-item Portrait Values Questionnaire-Revised (PVQ-RR). Confirmatory factor analysis (CFA) shows that a PCI-selected 10% subset substantially improves overall construct recovery relative to response-stability and random selection. These findings support PCI as a principled internal diagnostic for screening synthetic respondents in survey pipelines.