arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.06513cs.CLcs.LGstat.ML

SOL:通过双重切片Wasserstein度量衡量文本分布之间的差距

SOL: Measuring Gaps between Text Distributions by Double Sliced Wasserstein Metrics

  • TU Berlin(柏林工业大学)
  • University of Göttingen(哥廷根大学)

机构由 AI 辅助整理,请以论文原文为准。

Gregor Kornhardt, Moritz Piening, Jannis Chemseddine, Gabriele Steidl

AI总结:

提出SOL,一种基于双重切片Wasserstein距离的文本分布度量,用于评估非自回归模型,能检测分布失败并提供稳定估计。

AI中文摘要:

评估文本生成需要衡量生成分布与数据分布的匹配程度。对于自回归模型,这通过困惑度来实现。扩散和基于流的语言模型只能提供似然界,其紧密度在不同模型族之间有所差异。基于样本的替代方法,如带有熵的生成困惑度,未考虑分布拟合。我们提出SOL,一种文本分布之间的距离。每个序列通过固定变换器下其隐藏状态的经验测度来表示,并通过双重切片Wasserstein距离比较这些测度的分布。我们证明如果变换器是单射的,SOL是一种度量。实验表明,SOL能检测分布失败,恢复预期的模型趋势,并提供稳定的基于样本的估计。我们提出SOL以填补当前用于非自回归模型的评估协议中的空白。作为第一步,我们使用SOL重新评估在OpenWebText上训练的各种模型。

英文摘要:

Evaluating text generation requires measuring how well the generated distribution matches the data distribution. For autoregressive models, this is done by the perplexity. Diffusion and flow-based language models can only provide a likelihood bound, whose tightness differs between model families. Sample-based substitutes such as generative perplexity with entropy do not consider the distribution fit. We propose SOL, a distance between text distributions. Each sequence is represented by the empirical measure of its hidden states under a fixed transformer and the distributions of these measures are compared by the double sliced Wasserstein distance. We prove that SOL is a metric if the transformer is injective. Experiments show that SOL detects distributional failures, recovers expected model trends, and provides stable sample-based estimates. We put forward SOL to fill the gap in the current evaluation protocol used for non auto-regressive models. As a first step we use SOL to re-evaluate a variety of models trained on OpenWebText.

↑