arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25134cs.LGcs.AImath.PRstat.ML

大语言模型的概率结构

The Probabilistic Structure of Large Language Models

Adnan Aboulalaâ

AI总结:

本文以概率测度和随机过程统一描述大语言模型,将训练视为最大似然估计、生成视为序贯模拟,并探讨KL散度不对称性与幻觉的关系,同时类比扩散模型。

AI中文摘要:

本文从概率论视角审视大语言模型(LLMs),旨在以单一自洽的论述整合文献中通常分别处理的工具。LLMs通过令牌序列集合上的概率测度来描述,这些测度由其自回归条件分布指定。训练被表述为最大似然估计问题,通过随机梯度方法求解,而文本生成则被视为对所得随机过程的序贯模拟。本文考察了Kullback--Leibler散度的不对称性在文本生成中的作用,并将其与幻觉等典型现象以及统计合理性与真实性之间的区分联系起来。作为同一视角的补充说明,我们还讨论了围绕得分函数构建的扩散模型,该模型将生成视为逆时随机过程的模拟,该过程在离散和连续时间内将噪声转化为数据,而非序贯令牌预测。

英文摘要:

This paper presents a probabilistic perspective on large language models (LLMs), developed with the aim of bringing together, in a single self-contained account, tools that are usually treated separately across the literature. LLMs are described through probability measures on the set of sequences of tokens, specified via their autoregressive conditional distributions. Training is formulated as a maximum-likelihood estimation problem, addressed by stochastic gradient methods, while text generation is viewed as the sequential simulation of the resulting stochastic process. The role of the asymmetry of the Kullback--Leibler divergence in text generation is examined in relation with characteristic phenomena such as hallucination and the distinction between statistical plausibility and truth. As a complementary illustration of the same viewpoint, we also discuss diffusion models, built around the score function, which cast generation not as sequential token prediction but as the simulation of a reverse-time stochastic process transforming noise into data both in discrete and continuous time.

补充信息

↑