谱神经元
The Spectral Neuron
浏览论文内容
中文总结 AI 辅助
针对复杂模型丧失系数透明度与函数形状控制的问题,本研究提出谱神经元模型,通过读取仿射矩阵函数的特征值实现非线性预测,兼具可扩展性与数学可解释性,并验证了其实际训练与应用潜力。
中文摘要 AI 辅助
随着机器学习模型的复杂度和表达能力不断提升,简单模型的诸多特性(如内在系数透明度、对建模函数形状的控制能力)正逐渐丧失。在模型谱系的一端,是具备系数透明度但表达能力有限的简单线性模型;在谱系的另一端,则是表达能力随规模扩展而提升,但大多处于黑箱状态的神经网络。\n本研究提出了**谱神经元(spectral neuron)**概念:这是一种标量模型,表达式为 $f(x)=λ_k (A_0 + A_1 x + ... + A_n x_n)$,其中 $A_0, ..., A_n$ 为可学习的实对称矩阵。输入通过仿射矩阵函数进入模型,而预测结果通过读取该矩阵的一个特征值得到。因此,该模型虽为非线性模型,但其非线性来源在数学上仍是明确的。\n这为我们提供了一种实用的中间方案:随着矩阵维度增加,模型的表达能力可以不断增强,同时通过可学习的矩阵保留系数透明度。例如,极值特征值对应凸函数或凹函数,对系数矩阵施加半正定约束可实现单调性,而相关的特征空间可刻画局部特征的影响。\n我们研究了该模型族的系数透明度、特征影响边界和形状控制特性,随后测试了其在实际场景中能否被训练并实现规模扩展。我们对该模型族开展了系统性研究,结合多个数学领域的谱理论结果,刻画了其表达能力、系数透明度、特征影响和形状控制特性。相关代码可在 https://github.com/alexshtf/spectral_neuron_paper 获取。
英文摘要
As machine learned models increase in complexity and expressive power, features of simpler models, such as intrinsic coefficient transparency and control over the shape of the modeled function are lost. On the one edge of the spectrum we have simple linear models that possess coefficient transparency, but have a limited expressive power. On the other edge we have neural networks, that have expressive power that improves with scaling, but are mostly opaque. In this work we develop the \emph{spectral neuron} concept: a scalar model given by $f(x)=λ_k (A_0 + A_1 x + ... + A_n x_n)$, with learned real symmetric matrices $A_0, ..., A_n$. The input enters the model through an affine matrix function, but the prediction is obtained by reading one of its eigenvalues. Thus, the model is nonlinear, but the source of nonlinearity is still mathematically explicit. This gives us a useful middle ground: the model can become more expressive as the matrix dimension grows, while retaining coefficient transparency through the learned matrices. For example, extremal eigenvalues yield convex or concave functions, semidefinite constraints on the coefficient matrices impose monotonicity, and the associated eigenspaces characterize local feature influence. We study coefficient transparency, feature-influence bounds, and shape-control properties of this model family, and then test whether it can be learned and scaled in practice. We develop a systematic study of this model family, bringing together spectral results from several mathematical literatures to characterize its expressivity, coefficient transparency, feature influence, and shape-control properties. Code available at https://github.com/alexshtf/spectral_neuron_paper.
发表机构
- Technology Innovation Institute(技术创新研究院)
机构由 AI 辅助整理,请以论文原文为准。