发表机构
Hong Kong Polytechnic University(香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对仿射生成器与全连接 ReLU 架构,推导得出具有仿射隐参数化的神经网络的 Sharp 逼近率,明确了隐维度与网络预算的权衡关系,证明固定维度隐空间可实现误差随预算增大而消失。
AI 中文摘要
许多参数高效方法从低维隐表示生成大型神经网络的参数。给定具有 $P_\Phi$ 个参数槽的架构 $\Phi$,我们记 $\boldsymbol{\theta}_f=\mathcal{G}(\boldsymbol{\xi}_f)$,其中 $\mathcal{G}:\mathbb{R}^M\to\mathbb{R}^{P_\Phi}$ 是参数生成器,$\boldsymbol{\xi}_f\in\mathbb{R}^M$ 是目标函数 $f$ 的隐表示。架构 $\Phi$ 和生成器 $\mathcal{G}$ 在整个目标类中共享,而每个目标 $f$ 由其自身的隐向量 $\boldsymbol{\xi}_f$ 表示,$\Phi_{\mathcal{G}(\boldsymbol{\xi}_f)}$ 用于逼近 $f$。该框架涵盖超网络、低维参数化、参数高效适配和模型压缩。因此,理解隐维度 $M$ 和网络预算 $P$ 之间的权衡对于表征这些方法的表达效率至关重要。我们针对仿射生成器和全连接 ReLU 架构研究这一权衡。更准确地说,在满足 $P_\Phi\leq P$ 的架构 $\Phi$ 和仿射生成器 $\mathcal{G}:\mathbb{R}^M\to \mathbb{R}^{P_\Phi}$ 上联合优化,我们证明在 $[0,1]^d$ 上 $\alpha$-Hölder 函数单位球($0<\alpha\leq1$)上的最优最坏情况一致逼近误差具有 Sharp 阶 $\bigl(P\min\{M,P\}\bigr)^{-\alpha/d}$。特别地,我们的结果表明,即使固定维度的隐空间也足以在网络预算增加时实现逼近误差消失。
英文摘要
Many parameter-efficient methods generate the parameters of a large neural network from a low-dimensional latent representation. Given an architecture $Φ$ with $P_Φ$ parameter slots, we write $\boldsymbolθ_f=\mathcal{G}(\boldsymbolξ_f)$, where $\mathcal{G}\colon\mathbb{R}^M\to\mathbb{R}^{P_Φ}$ is a parameter generator and $\boldsymbolξ_f\in\mathbb{R}^M$ is a latent representation of the target function $f$. The architecture $Φ$ and the generator $\mathcal{G}$ are shared across the entire target class, while each target $f$ is represented by its own latent vector $\boldsymbolξ_f$, with $Φ_{\mathcal{G}(\boldsymbolξ_f)}$ approximating $f$. This framework encompasses hypernetworks, low-dimensional parameterizations, parameter-efficient adaptation, and model compression. Understanding the tradeoff between the latent dimension $M$ and the network budget $P$ is therefore fundamental to characterizing the expressive efficiency of these methods. We study this tradeoff for affine generators and fully connected ReLU architectures. More precisely, optimizing jointly over architectures $Φ$ satisfying $P_Φ\leq P$ and affine generators $\mathcal{G}:\mathbb{R}^M\to \mathbb{R}^{P_Φ}$, we prove that the optimal worst-case uniform approximation error over the unit ball of $α$-Hölder functions on $[0,1]^d$, where $0<α\leq1$, has the sharp order $ \bigl(P\min\{M,P\}\bigr)^{-α/d}. $ In particular, our result shows that even a fixed-dimensional latent space suffices to achieve vanishing approximation error as the network budget increases.