发表机构
AT&T(美国电话电报公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对概率电路曲率的组合特性,提出自适应锐度感知正则化,解决全局正则化的深度偏向与欠拟合问题,提升模型泛化能力。
AI 中文摘要
概率电路(Probabilistic Circuits, PCs)是支持精确推理的生成模型,与深度神经网络不同,它能给出精确且易处理的损失曲面曲率度量:对数似然的海森矩阵的迹。近期研究对该迹进行全局正则化,以引导学习过程偏向更平滑、泛化能力更强的最优解。我们表明,将锐度作为全局正则化项对PCs而言可能存在设定偏差,因为PCs的曲率本质上是可组合的。我们证明,每个求和节点对海森矩阵迹的贡献可精确分解为两部分:一是衡量该节点被使用程度的电路流,二是由其输出分布决定的局部锐度项。该分解解释了为何全局锐度正则化存在深度偏向,且可能导致欠拟合。基于此,我们引入一种自适应锐度感知正则化项,它根据节点的固有局部曲率对节点进行惩罚,并保留了闭式期望最大化(EM)更新。我们还通过实验表明,这种针对性正则化在保留锐度感知学习的鲁棒性和优势的同时,恢复了全局正则化所牺牲的泛化能力。
英文摘要
Probabilistic Circuits (PCs) are generative models that support exact inference and, unlike deep neural networks, admit an exact and tractable measure of loss-surface curvature: the trace of the Hessian of the log-likelihood. Recent work regularizes this trace globally to bias learning toward flatter, better generalizing optima. We show that treating sharpness as a global regularizer can be misspecified for PCs, whose curvature is inherently compositional. We prove that each sum node's contribution to the Hessian trace factorizes exactly into its circuit flow, which measures how heavily the node is used, and a local sharpness term determined by its output distribution. This decomposition provides insights into why global sharpness regularization is depth biased and can lead to underfitting. Building on it, we introduce an adaptive sharpness aware regularizer that penalizes nodes based on intrinsic local curvature and preserves closed form EM updates. We also show that empirically, this targeted regularization recovers the generalization that global regularization sacrifices while retaining the robustness and benefits of sharpness aware learning.