arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

单字普查:44种语言模型中的答案选择一致性

The One-Word Census: Answer-Choice Conformity Across 44 Language Models

Tapan Parikh

arXiv 2607.12796首次发表:更新:

发表机构

Cornell Tech(康奈尔科技)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究44种语言模型在单字选择上的一致性,用31个单轮提示词刻画,通过答案选择惊讶度评分,发现一致性程度高但模型间有差异且有结构,还对比了与人类规范差异,所有数据公开。

AI 中文摘要

当语言模型必须从大量同样有效的选项中选择一个答案时,它会选择哪个,其他模型选择相同答案的频率是多少?要求44个模型“选一个词,随便什么词”时,“serendipity”被选中的概率为41%。我们用31个单轮提示词来刻画这种一致性,每个提示词命名一个有多个有效单字答案的类别,每个模型无系统提示下问4次。分析基于标准化令牌的精确匹配,不使用嵌入和评判,每个模型成本约一美元。模型的一致性已有记录,我们的贡献在于普查工具本身及其揭示的一致性结构。我们通过答案选择惊讶度对每个模型评分,发现一致性程度很高,但模型间一致性差异超过四倍且有结构。角色和社区调整模型差异最大,最新主流旗舰模型最一致。在四个系列中,一致性随代际上升,但最新旗舰模型出现反转。排名对模型组成稳健。与人类类别生成规范相比,该领域在20个共享类别中的18个比人类更集中。所有提示词、记录和代码均公开。

英文摘要

When a language model must choose one answer from a large space of equally valid options, which answer does it choose, and how often is it the answer every other model chooses? Asked to "pick a word," 105 language models from more than twenty labs chose serendipity 46% of the time. We measure this convergence, and each model's share in it, with 96 single-turn prompts that each name a category with many valid one-word answers ("Name a tree."), asked eight times per model and scored by exact match, with no embeddings and no judge. A model's answer-choice surprisal is the average -log2 probability of its answers under the pooled answers of all other models. In 28 of 96 categories a single answer takes at least 80% of all answers. The concentration does not depend on small or persona-tuned models: the 87 major-lab models are at least as concentrated as the full field. Lightly post-trained and persona-tuned models are the most divergent; heavily post-trained assistants from the major labs are the most conformist. Models that avoid the modal answer mostly land on the same runner-up. Within the major providers' lineages, release order shows no panel-wide trend once model tier is controlled; GPT, Gemini, Grok and Qwen become more conformist across releases, and Claude's generation-5 releases reverse. On open checkpoints of three post-training pipelines, supervised fine-tuning is the largest step toward the field's answers. Against human category-production norms, the field is more concentrated than people in 18 of 20 shared categories. All prompts, transcripts, and code are public.

Commentsv3: 105 models, 96 prompts, 8 runs. Adds a major-lab vs small-model split, a spread-free score, training-stage checkpoints and a human-norms comparison. Data, code and transcripts: github.com/tap2k/modelun tag consensus-arxiv-v3)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑