检验而非假定充分性:针对涌现网络结构校准生成式社会模拟器
Testing, not presuming, adequacy: calibrating generative social simulators against emergent network structure
- Waseda University(早稻田大学)
- Gunma University(群马大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对生成式社会模拟器验证仅止于表面效度的问题,提出结合后验估计、充分性检验与诊断修复的校准协议,在二手奢侈品市场四单元中部分恢复充分性,并证明独立聚合模型无法再现全部网络结构特征。
AI中文摘要:
生成式社会模拟器的验证往往止步于表面效度:对涌现网络结构的比较是描述性的,既没有量化的参数不确定性,也没有充分性检验。我们提出了一种具有充分性意识的校准协议,该协议将摊销后验估计与合成可辨识性评估、匹配样本量的充分性检验(先验预测可达性加逐统计量的后验预测定位)、诊断引导的修复以及统计量留出审计相结合。我们在一个真实的二手奢侈品转售市场上进行了演示,该市场包含四个渠道×居住地单元,每个单元都是一个二分买家-品牌网络,使用由语言模型一次性离线提取的人物画像构建的前向模型。行为参数在全部四个单元中均可恢复,尽管校准是近似的,且对一个参数过度自信。观察到的汇总统计量在每个单元中都落在模拟器可达性参考之外,平均购买层级是普遍存在的差异。修复在四个单元中的两个满足了价值块标准,但未能恢复充分性,留出审计暴露了先前诊断未发现的买家广度离散度缺失。画像来源消融实验发现,语言模型画像在全部四个单元中均优于平坦规则基线,但类别内品牌重新标注未造成一致的退化,因此这些画像是一种部分验证的输入,其价值依赖于结构而非品牌身份。我们不提出因果主张,结论是:没有智能体交互或买家广度机制的独立聚合解释,无法共同再现市场的购买层级水平、头部品牌集中度、社区结构和买家广度异质性。
英文摘要:
Validation of generative social simulators often stops at face validity: emergent network structure is compared descriptively, without quantified parameter uncertainty or an adequacy check. We present an adequacy-aware calibration protocol that couples amortized posterior estimation with a synthetic identifiability assessment, a matched-sample-size adequacy check (prior-predictive reachability plus per-statistic posterior-predictive localization), a diagnosis-guided repair, and a statistic-held-out audit. We demonstrate it on a real second-hand luxury resale market with four channel-by-residency cells, each a bipartite buyer-brand network, using a forward model built from persona profiles elicited once, offline, by a language model. The behavioural parameters are recoverable in all four cells, though calibration is approximate and overconfident for one parameter. The observed summary falls outside the simulator's reachability reference in every cell, with the mean purchased tier as the pervasive discrepancy. The repair meets the value-block criterion in two of four cells but does not restore adequacy, and the held-out audit surfaces a buyer-breadth-dispersion miss no earlier diagnostic detected. A profile-source ablation finds the language-model profiles beat a flat rule baseline in all four cells, yet within-category brand relabelling causes no consistent degradation, so the profiles are a partially validated input whose value rests on structure, not brand identity. Making no causal claim, we conclude that an independent-aggregation account, without agent interaction or a buyer-breadth mechanism, cannot jointly reproduce the market's purchased-tier level, head-brand concentration, community structure and buyer-breadth heterogeneity.