什么决定投票胜负?法国 Compar:IA LLM 竞技场中的格式、长度与词汇多样性
What Wins a Vote? Formatting, Length, and Lexical Diversity in the French Compar:IA LLM Arena
- Compar:IA
- Université Paris Dauphine-PSL(巴黎多芬大学-PSL)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究通过文体计量学分析法国 Compar:IA 竞技场中的法语投票,发现粗体使用和词汇多样性(MATTR)与获胜关联最强,但调整排名未优于原始排名,建议并排发布以作敏感性分析。
AI中文摘要:
LLM 竞技场将成对的人类偏好转化为模型排名。这些偏好可能既反映答案的呈现方式,也反映其内容。我们对 2026 年 7 月 Compar:IA 发布版中的 137,293 条决定性法语投票采用了文体计量学方法;主要的格式分析涵盖 116 个模型的 137,113 场对战,联合估计则使用了具备所有必要测量的 127,092 场对战。对于每场对战,我们重建了用户投票时可见的回复。随后,我们将原始排名与针对格式、长度、可读性、词汇多样性和句子结构进行调整后的排名进行了比较。呈现方式与获胜相关,但长度、粗体文本和列表往往同时出现,使得它们各自的贡献难以分离。在所有测量特征中,两项关联在不同规格下变化最小:粗体使用(联合模型中每标准差胜率增加 +11.0%)和移动平均类符形符比(MATTR),这是一种对答案长度不太敏感的词汇多样性度量(+16.8%)。粗体关联在观察到的多轮对话中明显较小,而 MATTR 关联变化不大;由于用户选择是否继续对话,这种差异是描述性的而非因果性的。完整调整使 116 个模型中的 36 个排名至少移动了十位。然而,与外部基准的比较并未显示调整后的排名能更好地衡量能力。因此,我们建议将原始排名和调整后的排名并排发布,作为透明的敏感性分析。
英文摘要:
LLM arenas turn pairwise human preferences into model rankings. Those preferences may reflect how an answer is presented as well as what it says. We take a stylometric approach to 137,293 decisive French-language votes from the July 2026 Compar:IA release; the primary formatting analysis includes 137,113 battles across 116 models, and the joint estimates use the 127,092 battles with all required measurements. For each battle, we reconstruct the response visible when the user voted. We then compare the raw ranking with rankings adjusted for formatting, length, readability, vocabulary variety, and sentence structure. Presentation is associated with winning, but length, bold text, and lists tend to occur together, making their individual contributions hard to separate. Across the measured features, two associations change least across specifications: bold usage (+11.0% win odds per standard deviation in the joint model) and moving-average type-token ratio (MATTR), a measure of vocabulary variety that is less sensitive to answer length (+16.8%). The bold association is substantially smaller in observed multi-turn conversations, whereas the MATTR association changes little; because users choose whether to continue, this difference is descriptive rather than causal. The full adjustment moves 36 of 116 models by at least ten ranks. Yet comparisons with external benchmarks do not show that adjusted rankings better measure capability. We therefore recommend publishing raw and adjusted rankings side by side as a transparent sensitivity analysis.