发表机构
School of Computer and Communication Sciences, EPFL(计算机与通信科学学院,洛桑联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对布尔立方体上至多k次函数的高斯回归,改进了Polyanskiy-Samorodnitsky不确定性原理,得到了极小极大误差对应的尖锐采样阈值,揭示了噪声带来的指数级样本代价。
AI 中文摘要
我们在d维布尔立方体上的至多k次函数的已知m维子空间中,研究平方总体L₂损失下的高斯回归问题。随机输入可能对预测至关重要的区域采样不足,即使模型已知,也会延迟达到参数速率。对于固定的q₀<1/2、1≤k≤q₀d以及足够大的固定A,当置信度为1-e^{-t}(t≥log4)时,极小极大误差Aσ²(m+t)/n对应的最坏子空间样本阈值为N=(m+t)exp{E_{d,k}+O(k^{1/3})},其中E_{d,k}=dΨ(k/d),Ψ(q)=log2-H(1/2-√(q(1-q))),H为自然对数下的二元熵。该上界对所有可行的m成立;当m≤C(d,⌊k^{1/3}⌋)或t≥m时,匹配的下界成立。我们从两个方面改进了Polyanskiy-Samorodnitsky不确定性原理:其一,对于固定泄漏率ρ∈(0,1),承载非零至多k次多项式能量1-ρ比例的最小集合的概率为exp{-E_{d,k}+O_{ρ,q₀}(k^{1/3})},Airy核构造证明,一般情况下余项不能为o(k^{1/3});其二,我们构造了一个维度为C(d,⌊k^{1/3}⌋)的子空间,使得该子空间中的每个函数都在同一集合上拥有至少1-ρ比例的能量,该集合的概率至多为exp{-E_{d,k}+C_{ρ,q₀}k^{1/3}},当k足够大时,该集合为汉明球。一个显著的结果是噪声的指数代价:参数速率需要(m+t)4^k exp{-O(k^{1/3})}个样本,而无噪声识别仅需O((m+t)2^k)个样本;当k→∞且k/d→0时,有噪声阈值为(m+t)exp{2k+o(k)}。
英文摘要
We study Gaussian regression under squared population $L_2$ loss in a known $m$-dimensional subspace of degree-at-most-$k$ functions on the $d$-dimensional Boolean cube. Random inputs can undersample regions essential for prediction, delaying the parametric rate even when the model is known. For fixed $q_0<1/2$, $1\le k\le q_0d$, and sufficiently large fixed $A$, the worst-subspace sample threshold for minimax error $Aσ^2(m+t)/n$ with confidence $1-e^{-t}$, $t\ge\log4$, is \[ N=(m+t)\exp\{E_{d,k}+O(k^{1/3})\}, \quad E_{d,k}=dΨ(k/d), \] where $Ψ(q)=\log2-\mathsf H(\tfrac12-\sqrt{q(1-q)})$ and $\mathsf H$ is binary entropy with natural logarithms. The upper bound holds for every feasible $m$; the matching lower bound holds when $m\le\binom d{\lfloor k^{1/3}\rfloor}$ or $t\ge m$. We sharpen the Polyanskiy--Samorodnitsky uncertainty principle in two respects. First, for fixed leakage $ρ\in(0,1)$, the smallest set carrying a fraction $1-ρ$ of a nonzero degree-at-most-$k$ polynomial's energy has probability $\exp\{-E_{d,k}+O_{ρ,q_0}(k^{1/3})\}$. An Airy-kernel construction proves that the remainder cannot be $o(k^{1/3})$ in general. Second, we construct a subspace of dimension $\binom d{\lfloor k^{1/3}\rfloor}$ such that every function in the subspace has at least a fraction $1-ρ$ of its energy on the same set, whose probability is at most $\exp\{-E_{d,k}+C_{ρ,q_0}k^{1/3}\}$. For sufficiently large $k$, this set is a Hamming ball. A striking consequence is an exponential cost of noise: the parametric rate can require $(m+t)4^k\exp\{-O(k^{1/3})\}$ samples, whereas $O((m+t)2^k)$ suffice for noiseless identification. As $k\to\infty$ with $k/d\to0$, the noisy threshold is $(m+t)\exp\{2k+o(k)\}$.