AI 中文总结
该研究以四款中国前沿LLM智能体为对象,消除混淆因素后发现,合作倾向的设定单位是实验室,中国模型并非单一整体,且生态系统内差异大于东西差异。
AI 中文摘要
西方前沿大语言模型(LLM)智能体所记录的合作偏向是否延伸至不同对齐谱系?体现该谱系的中国模型应被视为单一整体,还是不同实验室的独立个体?本研究在进化型迭代囚徒困境中对四款前沿级中国模型——DeepSeek V4 Pro、Qwen3-Max、Kimi K2.5和GLM-5.1展开研究,采用的设计消除了先前研究中存在的混淆因素:未让各模型将自身自然语言策略转换为代码(该过程会使策略倾向与编码能力产生混淆),而是在所有实验室中固定使用转换器GPT-5.4 Mini,因此每一次跨实验室比较均为单纯的生成能力比较。本研究执行了完整实验方案:每条件下运行全对锦标赛与n=500的莫兰过程,涵盖三种提示风格与四种种群机制。研究评估了两项预先注册的假设:H6(非单一整体)得到支持:四款实验室在进攻性均衡比例P_A上存在显著差异,Qwen3-Max的P_A为1%,DeepSeek V4 Pro的P_A为9%,六组两两比较中有四组通过霍尔姆-邦费罗尼校正;四款实验室的差异(P_A范围8个百分点)大于中西方生态系统平均P_A的差异(均为5.0%),在此指标上,生态系统内差异超过东西差异。H5(合作偏向的普适性)虽一致但需限定:在12种实验室-提示组合中,有6种呈现合作多数,而西方模型的该比例为9/12,该差异不被视为确定结论,因为计数基于合作-中立的接近平局,且在预先注册的稳健性检验中使用替代转换器时,该比例升至9/12。研究结论为:合作倾向的设定单位是实验室,而非生态系统;证据不支持将“中国模型”视为单一整体的观点。
英文摘要
Does the cooperative bias documented for Western frontier LLM agents extend to a different alignment lineage, and should the Chinese models that embody it be treated as a single bloc or as distinct laboratories? We study four frontier-tier Chinese models - DeepSeek V4 Pro, Qwen3-Max, Kimi K2.5 and GLM-5.1 - in an evolutionary Iterated Prisoner's Dilemma, under a design that removes a confound present in prior work. Rather than letting each model convert its own natural-language strategies into code, which entangles strategic disposition with coding ability, we hold the converter fixed (GPT-5.4 Mini) across all labs, so every cross-lab comparison is a comparison of generation alone. We run the full protocol: all-play-all tournaments and a Moran process at n=500 runs per condition, across three prompt styles and four population regimes. Two pre-registered hypotheses are evaluated. H6 (not monolithic) is supported: the four labs differ significantly in aggressive-equilibrium proportion, P_A running from 1% for Qwen3-Max to 9% for DeepSeek V4 Pro, with four of six pairwise comparisons surviving Holm-Bonferroni. The spread across the four labs (P_A range 8pp) is larger than the difference between the Chinese and Western ecosystems' mean P_A (5.0% vs 5.0%): on this measure, within-ecosystem variation exceeds the East-West gap. H5 (cooperative-bias generality) is consistent but qualified: a cooperative plurality holds in 6 of 12 lab-prompt combinations against the 9 of 12 reported for Western models, a difference we do not treat as firm, since the count rests on Cooperative-Neutral near-ties and rises to 9/12 under an alternate converter in our pre-registered robustness check. The lab, not the ecosystem, is the unit at which cooperative disposition is set; treating "Chinese models" as a monolith is not supported by the evidence.
Comments9 pages, 8 tables. Companion study to arXiv:2605.29874, under a fixed-converter design. Code and replication package: https://github.com/arqFranciscoLeon/evollm (archived: https://doi.org/10.5281/zenodo.20248614)