大语言模型辩论在不同语言中重复论点的方式是否不同?
Do LLM Debates Repeat Arguments Differently Across Languages?
AI总结:
研究大语言模型辩论在不同语言中重复论点的差异,用“先验论点相似度”方法,发现在多语言嵌入模型中中文与英文有正向差距,该差距在多种条件下持续,多样性提示未显著缩小差距,建议多语言辩论评估衡量论证发展并报告缓解效果。
AI中文摘要:
大语言模型辩论通常通过最终答案进行评估,但文字记录也能揭示后续轮次是否产生了新的论证内容,或者是否用新的措辞回到了早期的主张。我们使用“先验论点相似度”来研究这一过程,这是一种综合诊断方法,用于比较在同一场辩论中提取的论点单元与早期单元。在针对71个议题、六种语言和四个模型智能体进行的八轮受控辩论中,中文是唯一一种在三种多语言嵌入模型中相对于英文始终存在正向差距的测试语言。该差距在不同智能体、轮次位置、回归调整、度量变体、提取长度控制、第二个提取器子集以及跨编码器尾部重新评分中均持续存在。人工校准显示项目级对齐较弱,但高相似度尾部富含实质性重复内容。一个具有多样性意识的提示会降低跨语言的先验论点相似度,但并未显著缩小中文与英文之间的差距。这些发现表明,多语言辩论评估应衡量随时间推移的论证发展,并以平均值和差距术语报告缓解效果。
英文摘要:
LLM debate is usually evaluated by final answers, yet transcripts reveal whether later turns develop new arguments or return to earlier claims in new wording. We study this process with \textit{prior-argument similarity}, which compares extracted argument units with earlier units in the same debate. In controlled eight-turn debates over 71 motions, six languages, and four model agents, Chinese is the only tested language with a consistently positive gap relative to English across three multilingual embedding models. The gap persists across agents, turn positions, regression adjustment, metric variants, extraction-length controls, a second-extractor subset, and cross-encoder tail rescoring. Manual calibration shows weak item-level alignment but a high-similarity tail enriched for substantive repetition. A diversity-aware prompt lowers \textit{prior-argument similarity} across languages, yet does not significantly narrow the Chinese--English gap. Multilingual debate evaluation should therefore measure argumentative development over time and report both average and gap terms.