arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.25292cs.AI

指令微调语言模型无法从它们能描述的分布中进行采样

Instruction-Tuned Language Models Cannot Sample from Distributions They Can Describe

发表机构韩国科学技术院计算学院 · 韩国科学技术院商业与技术管理学院
查看机构详情
  • School of Computing, KAIST(韩国科学技术院计算学院)
  • School of Business and Technology Management, KAIST(韩国科学技术院商业与技术管理学院)

机构由 AI 辅助整理,请以论文原文为准。

Chaemin Jang, Dongman Lee, Jihee Kim

首次发表
浏览论文内容

中文总结 AI 辅助

研究发现指令微调语言模型无法从其能描述的分布中采样,存在“知道/做”分裂,通过让模型描述分布可减半误差,还提出PPA在无额外成本下降低21%误差。

中文摘要 AI 辅助

硅采样将语言模型用作人类调查受访者的代理,将每个模型调用视为从人物角色的响应分布中独立抽取。我们表明这种抽取不存在:指令微调模型不会从分布中采样,它们会坍缩到单个输出。在一个民意基准测试中,同一个人物角色对同一个问题在超过一半的项目上返回相同答案。这种坍缩很明显:模型的内部概率集中在单个选项上,并且通过指令微调,失败情况大幅加剧。三个具有实质上不同训练后管道的模型家族中的每个指令微调模型在我们测试的每个任务上都失败,而基础模型失败的频率要低得多。引人注目的是,同一个无法从分布中采样的模型可以在一次调用中准确描述它。我们将这种差距称为“知道/做”分裂,并将其追溯到在对数its中可见且由对齐训练引起的退化采样原语。利用这种分裂,与人物角色聚合相比,让模型在一次调用中描述响应分布可将与人类调查数据相比的误差减半以上。对于需要按人物角色输出的应用程序,我们提出了Prompt - Perturbed Argyle(PPA),它在不增加成本的情况下将相同误差降低了21%。

英文摘要

Silicon sampling uses language models as proxies for human survey respondents, treating each model call as an independent draw from the persona's response distribution. We show this draw does not exist: instruction-tuned models do not sample from distributions, they collapse to a single output. The same persona on the same question returns the same answer on more than half of items in a public-opinion benchmark, and the model's internal probabilities concentrate on a single option. The failure is associated with, and amplified by, instruction-targeted post-training: instruction-tuned models are worse than their own bases in every family we can compare, the gap widens at each successive post-training stage and with the size of the tuning update, and continued pretraining on non-instruction tokens leaves it unchanged. Yet the knowledge survives: the same model that cannot sample from a distribution can describe it accurately in a single call. We call this gap the KNOWS/DOES split. Exploiting the split, a single call that asks the model to describe the response distribution more than halves the error against human survey data compared to persona aggregation. When per-persona outputs are required, we propose Prompt-Perturbed Argyle (PPA), which reduces the same error by 21\%, spreading each persona's answers to mirror real population differences at no added cost.

补充信息

↑