发表机构
Institute for Artificial Intelligence Research and Development of Serbia; University of Belgrade; Institute for Philosophy and Social Theory; Complexity Science Hub; University of Kragujevac; Faculty of Economics; University of Vienna; Faculty of Social Sciences; Faculty of Engineering; University of Novi Sad; Faculty of Philosophy; Nanyang Technological University; School of Social Sciences(塞尔维亚人工智能研发研究所; 贝尔格莱德大学; 哲学与社会理论研究所; 维也纳复杂性科学中心; 克拉古耶瓦茨大学; 经济学院; 维也纳大学; 社会科学学院; 工程学院; 诺维萨德大学; 哲学院; 南洋理工大学; 社会科学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过基准测试12种LLM配置,发现其在社交媒体反应预测中表现优异,可用于推荐系统压力测试,同时也揭示了大规模合成智能体群对公众意见的潜在威胁。
AI 中文摘要
社交媒体中的自主AI智能体对民主话语和平台治理构成切实风险,同时也为部署前推荐系统测试提供工具。一个核心开放性问题是,由人设提示的大语言模型(LLM)能否以足够的准确性模拟个体层面的社交媒体反应,以支持上述任一应用,以及准确性如何取决于个人资料完整性、模型选择和新帖子内容带来的泛化挑战。本研究在三种个人资料条件下,针对296个基于调查的智能体个人资料和26个经真实值映射的帖子,对12种LLM配置的二元点赞/踩预测进行基准测试,并采用留一帖子法的机器学习分类器作为基线。在完整个人资料条件下,准确率介于75.54%至96.68%之间,30个百分点的差异主要归因于模型选择,且经带智能体级自举区间的配对McNemar检验确认。GPT-5.5 Pro的准确率从完整个人资料下的96.68%,随个人资料精简降至62.32%,仅保留人口统计信息时进一步降至51.00%,后者与多数类基线无差异,证实人口统计推断提供的预测信号可忽略不计。留一帖子法下,监督分类器准确率骤降至15.4%,而LLM则能维持训练方法无法实现的真正零样本泛化。自适应推理可显著提升部分模型的准确率。对于带有直接个人资料锚点的帖子,模型间一致性(平均κ=0.44)几乎是无锚点帖子(κ=0.23)的两倍,且异质性最低的配置会使34%的模拟群体反应同质化。研究结果验证了基于LLM的模拟可用于推荐系统压力测试,同时表明其行为准确性使大规模合成智能体群对公众意见构成可信威胁。
英文摘要
Autonomous AI agents in social media present concrete risks to democratic discourse and platform governance, while also offering tools for pre-deployment recommender system testing. A central open question is whether persona-prompted LLMs can simulate individual-level social media reactions with sufficient accuracy to support either application, and how accuracy depends on profile completeness, model selection, and the generalization challenge posed by novel post content. This study benchmarks twelve LLM configurations on binary like/dislike prediction across 296 survey-based agent profiles and 26 ground-truth-mapped posts under three profile conditions, with leave-post-out machine learning classifiers as baselines. Across full-profile conditions, accuracy ranges from 75.54% to 96.68%, with a 30-point spread attributable primarily to model selection and confirmed by paired McNemar tests with agent-level bootstrap intervals. GPT-5.5 Pro accuracy degrades monotonically from 96.68% under a full profile to 62.32% under a reduced profile and to 51.00% with demographics alone, the last indistinguishable from the majority-class baseline, which confirms that demographic inference provides negligible predictive signal. Supervised classifiers collapse to 15.4% under leave-post-out, while LLMs sustain genuine zero-shot generalization unavailable to trained methods. Adaptive reasoning improves accuracy substantially for some models. Inter-model agreement is nearly double for posts with direct profile anchors (mean \k{appa} = 0.44) than for posts without them (\k{appa} = 0.23), and the least heterogeneous configuration homogenizes 34% of simulated population reactions. Results validate LLM-based simulation for recommender system stress-testing while documenting the behavioral accuracy that makes large-scale synthetic agent swarms a credible threat to public opinion.
Comments28 pages, 6 figures