arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

结构化输出会削弱44种语言模型的答案多样性

Structured Output Collapses Answer Diversity Across 44 Language Models

Tapan Parikh

arXiv 2607.18476首次发表:更新:

发表机构

Cornell Tech(康奈尔理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究语言模型在特定格式要求下答案多样性变化,通过对44个模型进行31个类别提示实验,发现结构化输出使答案收敛、多样性降低,存在寄存器梯度,揭示了模型对不同格式的响应差异及对答案多样性的影响。

AI 中文摘要

当语言模型必须从大量同样有效的选项中选择一个答案时,一个格式子句——“仅以JSON格式回复”——会改变它选择的答案。我们重新运行了单字普查(arXiv:2607.12796):向44个模型提出31个宽答案空间类别提示,现在要求以JSON格式回复——不进行模式强制,不进行约束解码,只有请求。收敛急剧加深:在无约束的“选一个词”提示中,模态答案从占库的41%升至64%,不同答案从52个降至36个;平均答案选择惊奇度从1.80比特降至1.58比特。这种影响是渐进的:44个模型中有6个单独变动(BH-FDR q =.10),都朝着模态方向,由最具特色的模型引领,而顺从底线不变。它是一个锐化器,而非重新索引器——31个类别中有28个类别中的普通聊天模态答案得以保留。默认设置按寄存器索引:在一次运行内重新采样(n = 20)发现JSON改变了模型53%的稳定聊天默认设置,大多变回大众答案,并设置了聊天中不存在的默认设置(Claude Fable对颜色0%的时间在聊天中回答“天蓝色”,在JSON中为100%)。全电池控制揭示了一个寄存器梯度:压缩显著且特定于模型训练所使用的答案交付格式(JSON -0.22比特,p =.0002;XML -0.19,p =.002),YAML和CSV不存在,对任意括号包装则相反(+0.13,p =.009)——这将机制倾向于训练后工具使用。在解码器处强制模式(响应格式)的压缩不超过请求(-0.03比特):这种坍缩存在于模型对寄存器的响应中,而非解码器。结构化输出是软件使用语言模型的方式,与用于评估、比较和选择模型的聊天表面相比,该表面由一个明显更同质的模型提供服务。

英文摘要

When a language model must choose one answer from a large space of equally valid options, a format clause -- "Reply with JSON only" -- changes which answer it chooses. We re-run the One-Word Census (arXiv:2607.12796): 31 wide-answer-space category prompts asked of 44 models, now with the reply requested in JSON -- no schema enforcement, no constrained decoding, only the request. Convergence deepens sharply: on the unconstrained "Pick a word" prompt the modal answer rises from 41% to 64% of the pool and distinct answers fall from 52 to 36; mean answer-choice surprisal drops from 1.80 to 1.58 bits. The tax is progressive: six of 44 models move individually (BH-FDR q=.10), all toward the mode, led by the most distinctive models, while the conformist floor is immobile. It is a sharpener, not a re-indexer -- the plain-chat modal answer survives in 28 of 31 categories. Defaults are register-indexed: a within-run re-sample (n=20) finds JSON shifts 53% of a model's stable chat defaults, mostly back to the crowd, and installs defaults absent from chat (Claude Fable 5 answers "cerulean" for colour 0% of the time in chat, 100% in JSON). Full-battery controls reveal a register gradient: compression is significant and specific to the answer-delivery formats models are trained to speak (JSON -0.22 bits, p=.0002; XML -0.19, p=.002), absent for YAML and CSV, and reversed for an arbitrary bracket wrapper (+0.13, p=.009) -- weighing the mechanism toward tool-use post-training. Enforcing the schema at the decoder (response_format) compresses no further than the request (-0.03 bits): the collapse lives in the model's response to the register, not the decoder. Structured output is how software consumes language models, and that surface is served by a measurably more homogeneous model than the chat surface on which models are evaluated, compared, and chosen.

Comments12 pages, 1 figure. Companion to the One-Word Census (arXiv:2607.12796). Code, data, and interactive explorer: https://github.com/tap2k/modelun/tree/main/studies/structured

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑