arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

乘积Wasserstein流形上学习的统计力学

Statistical Mechanics of Learning on Product Wasserstein Manifolds

Srinivasa Rao P Vangmayi P Reddy

arXiv 2608.01434首次发表:更新:

发表机构

Curlvee Technolabs; C-DAC; Indian Institute of Technology Madras(库尔维科技公司; 印度高级计算发展中心; 印度马德拉斯理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究将结构约束重定义为几何先验,提出在乘积Wasserstein流形上开展学习,开发相关算法,实验表明其可提升学习系统的泛化能力、稳定性并缓解贫瘠高原问题。

AI 中文摘要

通常,学习的统计力学将权重分布上的约束视为缩小可能解空间的限制,因此会降低模型容量。在本文中,我们希望采取一种相反的方法,该方法基于先前关于分布约束感知机的研究工作。我们不将规定的权重分布视为单纯的限制,而是提出它定义了学习自然展开的内在几何结构。我们将深度神经网络和变分量子电路均公式化为Wasserstein流形乘积上的梯度流——每层对应一个经典Wasserstein空间,电路参数对应一个量子Wasserstein空间。在该几何结构中,先前与分布约束相关的容量降低,表现为约束流形本身的度量结构。我们为深度网络开发了分层平均场描述,使用1阶量子Wasserstein距离将框架扩展到量子场景,并引入两种实用算法:分层DisCo-SGD和Quantum DisCo,它们遵循乘积流形上的近似测地线。在师生问题、标准图像分类任务以及小型变分量子分类器上的实验表明,与无约束和纯基于范数的基线相比,尊重这些分布几何结构可提升泛化能力、稳定训练并减少贫瘠高原的严重程度。该方法首次将结构约束重新定义为几何先验,为将生物学、光谱学或硬件衍生的分布信息融入经典和量子学习系统提供了途径。

英文摘要

Normally the statistical mechanics of learning treats constraints on weight distributions as restrictions that shrink the space of possible solutions. Therefore, it reduces model capacity. In this paper we would like to take a contrary approach, which, however, is based on the earlier work on distribution-constrained perceptrons. Rather than treating a prescribed weight distribution as a mere restriction, we propose that it defines the intrinsic geometry upon which learning naturally unfolds. We formulate both deep neural networks and variational quantum circuits as gradient flows on a product of Wasserstein manifolds -- one classical Wasserstein space for each layer and one quantum Wasserstein space for the circuit parameters. Within this geometry, the capacity reduction, which was previously associated with distributional constraints, appears as the metric structure of the constraint manifold itself. We develop a hierarchical mean-field description for deep networks, extend the framework to the quantum setting using the quantum Wasserstein distance of order 1, and introduce two such practical algorithms, Hierarchical DisCo-SGD and Quantum DisCo, that follow approximate geodesics on the manifold of the product itself. Experiments on teacher-student problems, standard image classification tasks, and small variational quantum classifiers show that respecting these distributional geometries improves generalization, stabilizes training, and reduces the severity of barren plateaus compared with unconstrained and purely norm-based baselines. This approach firstly reframes structural constraints as geometric priors and suggests a route for incorporating biological, spectral, or hardware-derived distributional information into both learning systems, viz., classical and quantum learning.

Comments18 pages, 5 Figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑