发表机构
University College London; Holistic AI; University of Utah(伦敦大学学院; 整体人工智能公司; 犹他大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文阐述Tsallis统计在AI领域的应用,梳理其数学核心,综述其在多类AI任务中的应用模式,指出深度网络相关统计特征的非广延性,主张将$q$作为可学习归纳偏置。
AI 中文摘要
Tsallis统计通过单个实参数$q$推广了玻尔兹曼-吉布斯统计力学,该参数控制对稀有和频繁事件分配的权重。该框架最初用于描述具有长程关联、多重分形几何和重尾波动的物理系统,现已成为现代人工智能(AI)的重要组成部分:它是稀疏注意力机制(\textsc{sparsemax}和$\textsc{α-entmax}$)、具有可控探索的最大熵强化学习、鲁棒重尾概率模型以及一系列广义损失函数和正则化项的基础。本文提供了Tsallis统计与AI交叉领域的结构化视角。我们首先回顾数学核心:$q$-熵及其变分(最大熵)基础、$q$-指数和$q$-对数、$q$-中心极限定理、$q$-高斯分布及其在超统计中的动力学起源,重点关注对机器学习重要的性质。随后我们综述其在softmax泛化、强化学习、序列和图神经模型、生成式和概率建模、损失设计及优化中的应用,提取每种情况下的重复设计模式:由$q$控制的密集/均匀与稀疏/峰值行为之间的可调插值。我们进一步指出,深度网络中经验观察到的重尾权重谱和梯度噪声统计本身就是非广延性特征,将现代学习动力学置于$q$-统计的范围内。最后,我们讨论方法学陷阱、与信息几何和$q$-指数族的关系以及开放方向,主张应将$q$视为可学习的归纳偏置而非固定超参数。
英文摘要
Tsallis statistics generalizes Boltzmann-Gibbs statistical mechanics through a single real parameter $q$ that controls the weight assigned to rare and frequent events. Originally proposed to describe physical systems with long-range correlations, multifractal geometry, and heavy-tailed fluctuations, the framework has become a recurring ingredient in modern artificial intelligence (AI): it underlies sparse attention mechanisms (\textsc{sparsemax} and $α$-\textsc{entmax}), maximum-entropy reinforcement learning with controllable exploration, robust and heavy-tailed probabilistic models, and a family of generalized loss functions and regularizers. This paper offers a structured perspective on where Tsallis statistics meets AI. We first review the mathematical core: $q$-entropy and its variational (maximum-entropy) foundation, the $q$-exponential and $q$-logarithm, the $q$-central limit theorem, $q$-Gaussian distributions, and their dynamical origin in superstatistics, emphasizing the properties that matter for machine learning. We then survey applications across softmax generalization, reinforcement learning, sequential and graph neural models, generative and probabilistic modeling, loss design, and optimization, extracting the recurring design pattern in each case: a tunable interpolation between dense/uniform and sparse/peaked behavior governed by $q$. We further argue that the heavy-tailed weight spectra and gradient-noise statistics empirically observed in deep networks are themselves nonextensive signatures, placing modern learning dynamics within the scope of $q$-statistics. Finally, we discuss methodological pitfalls, the relationship to information geometry and $q$-exponential families, and open directions, arguing that $q$ should be treated as a learnable inductive bias rather than a fixed hyperparameter.