发表机构
IIT Kanpur; KU Leuven(印度理工学院坎普尔分校; KU Leuven(比利时鲁汶大学))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出结合深度生成模型与符号回归的框架,从样本中估计非归一化分布的符号密度,在多元玩具分布、XY模型等任务中有效,可推导φ⁴理论重正化问题的哈密顿函数近似。
AI 中文摘要
从观测样本中估计概率密度函数(PDF)的符号或解析形式,是统计与计算建模中的基础挑战,该过程对推导表征潜在现象的可解释、可泛化关系至关重要。传统上,这种估计高度依赖领域专业知识与特定领域先验知识,专家需基于经验证据和理论理解选择合适的函数形式或参数族,再通过参数估计确定这些形式的系数。本文提出一种框架,利用领域特定先验知识(如相互作用范围和预定义的基元函数)从观测样本中估计非归一化分布的符号表达式,将深度生成模型与符号回归(SR)相结合,融入如分解大型分布等归纳偏置以保持问题可处理性,所研究的深度生成模型包括基于似然的流模型和基于分数的模型。实验表明,该框架在估计多元玩具分布及计算物理中的格点(即XY模型和φ⁴理论)的密度函数时具有有效性;当应用于φ⁴理论的重正化问题时,该框架可直接从样本中估计不同尺度下哈密顿函数的紧凑符号近似,得到的表达式在非微扰场景下可能难以用传统微扰或解析方法推导。
英文摘要
Estimating the symbolic or analytical form of probability density functions (PDFs) from observed samples is a fundamental challenge in statistical and computational modelling. This process is critical for deriving interpretable and generalizable relationships characterizing the underlying phenomenon. Traditionally, this estimation depends strongly on domain expertise and prior field-specific knowledge, with experts selecting appropriate functional forms or parametric families based on empirical evidence and theoretical understanding. The coefficients of these forms are then typically determined through parameter estimation. In this paper, we develop a framework for estimating symbolic expressions of unnormalized distributions from observed samples using domain-specific prior knowledge, such as the range of interactions and a predefined set of primitive functions. We integrate deep generative models with symbolic regression (SR), incorporating inductive biases, such as factorizing large distributions, to keep the problem tractable. The deep generative models we examine include likelihood-based models, viz., flow models, and score-based models. Experiments show the effectiveness of the proposed framework for estimating density functions for multivariate toy distributions as well as lattices from computational physics, namely, XY model and $ϕ^4$ theory. When applied to the renormalization problem in $ϕ^4$ theory, the proposed framework estimates compact symbolic approximations of the hamiltonian function at different scales directly from samples, yielding expressions that may be challenging to derive using traditional perturbative or analytic approaches in nonperturbative settings.
Comments32 pages, 3 figures
Journal refTransactions on Machine Learning Research, issn={2835-8856},year={2026}