arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型是否捕捉到了其训练数据中的多样性?

Do Large Language Models Capture the Diversity in their Training Data?

Youqi Wu, Farzan Farnia

arXiv 2609.02275首次发表:更新:

发表机构

The Chinese University of Hong Kong(香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究从信息论视角发现大语言模型等生成模型存在系统性条件多样性差距,提出事后修正机制并证明相关凹性以缓解该差距。

AI 中文摘要

大语言模型被训练以对文本的条件分布进行建模,但人们对它们是否捕捉到了训练数据中存在的所有合理输出的多样性仍了解不足。我们从信息论视角研究这一问题,将模型生成输出的条件熵与对应训练数据的条件熵进行比较。给定配对的输入-输出样本,我们利用条件熵及其基于冯·诺依曼熵的矩阵类似物,来测量条件输入未解释的输出变异性,无需同一提示的多个参考输出。在拥有公开训练数据的LLM系列中,包括OLMo、Pythia和GPT-Neo,我们在不同模型规模、序列长度和解码策略下,始终发现模型生成输出的条件熵低于其训练数据。我们在语言建模之外也观察到类似的条件多样性差距,包括类条件ImageNet生成器和在MS-COCO上训练的文本条件模型。为解决这一差距,我们提出一种事后修正机制,为每个输入生成多个输出并通过矩阵熵投影对其重新加权,在保持接近原始模型分布的同时增加条件多样性。我们证明了基于矩阵的条件熵泛函的凹性,这使得所得的熵约束投影成为凸优化问题,并开发了可扩展的镜像下降算法用于其实现。我们的结果揭示了现代生成模型与其训练数据之间存在系统性的条件多样性差距,并提供了用于测量和缓解该差距的信息论框架。

英文摘要

Large language models are trained to model conditional distributions over text, yet it remains inadequately understood whether they capture the full diversity of plausible outputs present in their training data. We study this question through an information-theoretic lens by comparing the conditional entropy of model-generated outputs with that of the corresponding training data. Given paired input-output samples, we use conditional entropy and its matrix-based analogue based on von Neumann entropy to measure output variability beyond what is explained by the conditioning input, without requiring multiple reference outputs for the same prompt. Across LLM families with publicly available training data, including OLMo, Pythia, and GPT-Neo, we consistently find that model-generated outputs exhibit lower conditional entropy than their training data, across different model scales, sequence lengths, and decoding strategies. We observe a similar conditional diversity gap beyond language modeling, including class-conditioned ImageNet generators and text-conditioned models trained on MS-COCO. To address this gap, we propose a post-hoc correction mechanism that generates multiple outputs for each input and reweights them through a matrix-entropy projection, increasing conditional diversity while remaining close to the original model distribution. We prove the concavity of the matrix-based conditional entropy functional, which makes the resulting entropy-constrained projection a convex optimization problem, and develop a scalable mirror-descent algorithm for its implementation. Our results reveal a systematic conditional diversity gap between modern generative models and their training data, and provide an information-theoretic framework for measuring and mitigating this gap.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑