发表机构
The University of Hong Kong; Sun Yat-sen University(香港大学; 中山大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究深度神经网络通过函数复合实现逼近的机制,证明分段线性生成器表示的刚性限制,并构造光滑生成器实现双重指数误差衰减,为深度分配提供理论指导。
AI 中文摘要
深度神经网络通过将仿射映射与非线性激活函数复合来逼近函数,但复合本身如何产生逼近能力尚未被完全理解。我们研究了一个基本机制:单个标量生成函数迭代的几何加权和。该机制支撑了函数\(x - x^2\)的经典帐篷映射构造,以及Yarotsky、W. E等人用于分析深度神经网络逼近能力的相关递归表示。首先,我们建立了一个刚性定理:对于具有有限段数的连续分段线性生成器,任何能以这种方式表示的\(C^3\)函数至多是二次的。对于非仿射二次函数,几何因子至少为$1/4$。这一结果既揭示了帐篷映射方法的局限性,又补充了基于分层基和递归多项式构造的现有方法。其次,以精确余项恒等式为指导,我们构造了一个光滑生成器,其迭代在平方逼近中产生关于总深度的双重指数误差衰减,并通过乘法模块对每个固定多项式也是如此。对于在\([-1,1]^d\)上具有绝对可求和系数的幂级数,根据单项式次数分配深度,在每个内部立方体上产生\(O(e^{-cL^{1/d}})\)阶的一致逼近误差。这些发现表明生成器动力学和余项估计如何控制深度神经网络的深度分配和逼近速率。
英文摘要
Deep neural networks approximate functions by composing affine maps with nonlinear activations, but how composition itself creates approximation power is not yet fully understood. We investigate a fundamental mechanism: geometrically weighted sums of iterates of a single scalar generator function. This mechanism underpins the classical tent-map construction of the function \(x - x^2\) and related recursive representations used by Yarotsky, W. E, et al., to analyze the approximation powers of deep neural networks. First, we establish a rigidity theorem: for continuous piecewise linear generators with a finite number of segments, any \(C^3\) function that can be represented in this way is at most quadratic. For non-affine quadratic functions, the geometric factor is at least $1/4$. This result both reveals limitations of the tent-map approach and complements existing methods based on hierarchical bases and recursive polynomial constructions. Second, using an exact remainder identity as guidance, we construct a smooth generator whose iterates yield doubly exponential error decay in total depth for square approximation and, through multiplication modules, for each fixed polynomial. For power series with absolutely summable coefficients on \([-1,1]^d\), distributing depth according to monomial degree yields a uniform approximation error of order \(O(e^{-cL^{1/d}})\) on each interior cube. These findings demonstrate how generator dynamics and remainder estimates govern depth allocation and approximation rates of deep neural networks.