arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16454cs.AI

微调修复大语言模型中的模式坍缩与过度离散

Fine-Tuning Fixes Mode Collapse and Over-Dispersion in LLMs

Kirill Skobelev, Eric Fithian, X. Y. Han

首次发表
浏览论文内容

中文总结 AI 辅助

本文证明大语言模型输出的模式坍缩或过度离散取决于模型与数据,且足够的监督微调可使多样性收敛于目标分布,并通过偏差-方差分解与KL散度界限从理论上刻画了这一过程。

中文摘要 AI 辅助

Doshi和Hauser(2024)、Bisbee等人(2024)以及Xie等人(2026)的近期工作提出担忧,认为大语言模型(LLMs)的输出往往多样性不足:它们重复或彼此相似的程度高于其旨在代表的人群中的回答,这种现象被称为模式坍缩。在本工作中,我们表明是否发生模式坍缩或其相反情况取决于所使用的具体模型和数据集。此外,在足够的监督微调(SFT)数据下,LLM输出的多样性收敛于微调数据所采样的目标分布的多样性。为量化这一比较,我们测量从同一固定提示中独立采样的两个响应重合(碰撞)的概率,或它们在核函数下的期望相似度。我们推导了模型与目标碰撞概率之间期望差距的偏差-方差分解,表明SFT并非固有地偏向模式坍缩或其相反情况:有限样本的SFT可能使模型处于欠离散或过度离散状态,具体取决于模型和数据集。最后,我们表明绝对差距受从目标分布到模型的Kullback-Leibler(KL)散度的平方根约束。因此,在总体交叉熵下足够接近最优的模型不可能表现出任意失准的多样性。我们在三个实验中测试了该分解和界限:在合成语言上的小型Transformer、在人类调查上微调的四个LLM,以及这些LLM在CodeNet(人类代码解决方案数据集)上微调。在所有实验中,更多的目标数据使模型多样性向人类(或合成目标)水平移动,与我们的理论预测一致。这些结果表明,多样性失准可能源于有限样本误差,并随着SFT更好地逼近目标分布而缩小。

英文摘要

Recent work by Doshi and Hauser (2024), Bisbee et al. (2024), and Xie et al. (2026) raises concerns that outputs from large language models (LLMs) tend to be under-diverse: they repeat or resemble one another more often than responses from the population they are meant to represent, a phenomenon known as mode collapse. In this work, we show that whether mode-collapse, or its opposite, occurs depends on the specific model and dataset used. Further, with sufficient supervised fine-tuning (SFT) data, LLM output diversity converges toward that of the target distribution from which fine-tuning data are sampled. To quantify this comparison, we measure the probability that two responses sampled independently from the same fixed prompt coincide (collide), or their expected similarity under a kernel. We derive a bias-variance decomposition of the expected gap between the model's and target's collision probabilities, showing that SFT is not inherently biased toward mode collapse or its opposite: finite-sample SFT can leave a model either under- or over-dispersed, depending on the model and dataset. Finally, we show that the absolute gap is bounded by the square root of the Kullback-Leibler (KL) divergence from the target distribution to the model. Consequently, a model sufficiently close to optimal under population cross-entropy cannot exhibit arbitrarily miscalibrated diversity. We test the decomposition and the bound in three experiments: small transformers on synthetic languages, four LLMs fine-tuned on human surveys, and these LLMs fine-tuned on CodeNet, a dataset of human code solutions. More target data moves model diversity toward the human (or synthetic target) level in all experiments, consistent with our theoretical predictions. These results show that diversity miscalibration can arise from finite-sample error and shrink as SFT better approximates the target distribution.

发表机构

  • Northwestern University(西北大学)
  • University of Chicago(芝加哥大学)

机构由 AI 辅助整理,请以论文原文为准。

↑