AI 中文总结
该研究开发了校准后的SimSoM社交媒体模拟器,通过实验验证后发现静态审核会高估用户封禁效果,实时审核时补偿性转发会削弱低质量内容减少效果,为内容审核政策评估提供了基于模拟的框架。
AI 中文摘要
基于智能体的社交媒体模拟器为研究内容审核提供了可控环境,但其价值取决于它们重现真实平台动态的忠实程度。我们开发了SimSoM的校准扩展版本,这是一个基于智能体的社交网络信息扩散模型,以COVID-19大流行期间在线疫苗讨论的真实世界数据集为基础。我们的方法用经验拟合的分布替代了特设参数化,通过CMA-ES(协方差矩阵自适应进化策略)进行优化,并在时间、分布和结构维度上与真实数据进行了验证。使用这个经过验证的模拟器,我们提供了三个关键贡献:第一,我们表明校准后的模型重现了经验数据的关键统计特征,包括活动分布、帖子/转发比率和时间模式;第二,我们将已建立的错误信息传播者检测和预防方法应用于经验数据和模拟数据,逐步移除排名靠前的用户,结果显示低质量内容的减少在两种数据中是一致的;第三,在30次网络实现中比较静态(追溯性)和动态(模拟内)审核,我们表明静态评估显著高估了最有效检测器的用户封禁效果:当实时应用审核时,剩余用户的补偿性转发会抑制低质量内容的预期减少,因此静态估计应被视为上限。这些发现强调了基于模拟的内容审核政策评估的必要性,并提供了一个可重复使用、基于经验的模拟框架。
英文摘要
Agent-based social media simulators offer a controlled environment to study content moderation, yet their value hinges on how faithfully they reproduce real platform dynamics. We develop a calibrated extension of SimSoM, an agent-based model of information diffusion on social networks, grounded in a real-world dataset of online vaccine discourse during the COVID-19 pandemic. Our approach replaces ad-hoc parametrisations with empirically fitted distributions, optimised via CMA-ES (Covariance Matrix Adaptation Evolution Strategy) and validated against real data across temporal, distributional, and structural dimensions. Using this validated simulator, we provide three key contributions. First, we show that the calibrated model reproduces key statistical signatures of the empirical data, including activity distributions, post/reshare ratios, and temporal patterns. Second, we apply established misinformation-spreader detection and prevention methods to both empirical and simulated data, progressively removing top-ranked users and showing that the resulting decline in low-quality content is consistent across the two. Third, comparing static (retroactive) and dynamic (in-simulation) moderation across 30 network realisations, we show that static evaluation significantly overestimates the effectiveness of user bans for the most effective detectors: when moderation is applied in real time, compensatory resharing by the remaining users dampens the expected reduction in low-quality content, so static estimates should be read as an upper bound. These findings highlight the necessity of simulation-based evaluation for content moderation policies and contribute a reusable, empirically grounded simulation framework.
Comments15 pages, 4 figures, 2 tables. Accepted for presentation at the Social Simulation Conference (SSC) 2026