AI 中文总结
本研究通过PV-SST测试床开展实验,发现同行排名feed会诱导LLM智能体群体的词汇趋同,但未证实分布式来源相比单一来源具有可靠的立场改变优势。
AI 中文摘要
大语言模型(LLM)智能体的群体行为无法通过单智能体基准来表征。我们引入PV-SST(peer-voted social-platform testbed,同行投票社交平台测试床),并报告一项预先注册的匹配暴露实验,该实验涵盖四个主题、四个未使用的随机种子、四个开放权重模型家族以及三个预先指定的更大变体。实验包含448次试验和112个完整的模型-主题-种子块。与仅主题对照组相比,按同行生成点赞数排序的前一轮同行帖子feed,在四个家族核心面板中使最终轮的词汇相似度增加(配对均值差为+0.0082 TF-IDF余弦单位,95%块自助法置信区间[0.0043, 0.0121],随机化p值=0.000105,n=64个块),在三个变体规模扩展中同样增加(+0.0109 [0.0069, 0.0151],p=0.000001,n=48)。这种对比将同行帖子暴露与排名绑定,因此无法识别仅排名的效应。核心面板中对立侧的存活率下降了3.9个百分点([-6.8, -1.6],p=0.0068),但在更大变体中无明确结论(-1.0个百分点[-3.1, 0.4],p=0.50)。在保持对抗性印象固定的情况下,四个分布式来源并不比单个来源更可靠地改变诚实智能体的立场。预先注册的分布式-单一对比在核心面板中为正但无明确结论(+0.057 [-0.009, 0.125],p=0.112),在更大变体中为负(-0.040 [-0.113, 0.035],p=0.332),未通过预先指定的跨模型和跨主题一致性标准。因此,稳健的结果是测试的同行排名feed下的词汇趋同,而非普遍的观点捕捉或普遍的协调优势。该研究评估合成LLM智能体群体,不估计对人类或生产平台的影响。
英文摘要
Population-level behavior in large-language-model (LLM) agents cannot be characterized by single-agent benchmarks. We introduce PV-SST, a peer-voted social-platform testbed, and report a separately frozen, preregistered matched-exposure experiment spanning four topics, four unused seeds, four open-weight model families, and three prespecified larger variants. The experiment comprises 448 trials and 112 complete model-by-topic-by-seed blocks. Relative to a topic-only control, a feed of previous-round peer posts ranked by peer-generated likes increases final-round lexical similarity in both the four-family core panel (paired mean difference +0.0082 TF-IDF cosine units, 95% block-bootstrap CI [0.0043, 0.0121], randomization p=0.000105, n=64 blocks) and the three-variant size extension (+0.0109 [0.0069, 0.0151], p=0.000001, n=48). This contrast bundles peer-post exposure with ranking and therefore does not identify a ranking-only effect. Opposite-side survival falls in the core panel (-3.9 percentage points [-6.8, -1.6], p=0.0068) but not conclusively in the larger variants (-1.0 pp [-3.1, 0.4], p=0.50). Holding adversarial impressions fixed, four distributed sources do not reliably move honest-agent stance more than one source. The preregistered distributed-minus-single contrast is positive but inconclusive in the core panel (+0.057 [-0.009, 0.125], p=0.112) and negative in the larger variants (-0.040 [-0.113, 0.035], p=0.332), failing the prespecified cross-model and cross-topic consistency criterion. Thus the robust result is lexical convergence under the tested peer-ranked feed, not general opinion capture or a general coordination advantage. The study evaluates synthetic LLM-agent populations; it does not estimate effects on people or production platforms.
Comments9 pages, 3 figures. Code, frozen protocol, configurations, summary tables, and data are publicly available at https://github.com/ranausmanai/synthetic-social-networks and https://huggingface.co/datasets/ranausmans/synthetic-social-networks