发表机构
University of Minnesota(明尼苏达大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文证明在固定参数预算下,循环估计器通过共享参数重复计算,可在不增加参数的情况下提升统计精度,并在多种模型上达到极小极大最优速率。
AI 中文摘要
人工智能中不断增长的存储需求促使我们使用更少的可训练参数进行学习。我们研究在共同的参数预算下,循环估计器(即重复应用一个在各迭代间共享参数的拟合算子)能否提高统计精度。其传统的非绑定对应方法在每次迭代中使用独立的参数。对于一般似然模型,我们为循环筛极大似然估计建立了平方Hellinger风险的上界,并为调优的非绑定族建立了极小极大下界。这些界揭示了一个参数-迭代-精度权衡:重复计算可以在不增加参数的情况下改善近似,同时增加计算成本和拟合类的复杂度。对于已知Hölder光滑性的目标,循环残差前馈网络和指定的层归一化后Transformer在固定数量的有界实数参数下,达到极小极大多项式速率(至多相差对数因子)。在足够大的固定预算下,循环最坏情况风险随样本量增长而趋于零,而最优非绑定最坏情况风险仍远离零。在指定的增长预算条件下,循环与非绑定风险之比也趋于零。高斯和拉普拉斯回归、二元响应以及基于能量的密度估计验证了该理论。
英文摘要
Memory constraints in artificial intelligence motivate accurate function approximation with fewer parameters. We study looping, which repeatedly composes one update function with shared parameters; each output becomes the next input. A looped Transformer, for example, reuses one block, whereas its conventional untied counterpart uses separately parameterized blocks. We compare their parameter requirements for a given worst-case approximation accuracy, or equivalently, their approximation accuracy under a common budget limiting distinct trainable coefficients. We then ask whether this representational parsimony improves statistical accuracy. For general likelihood models, we establish an upper squared Hellinger risk bound for looped sieve maximum likelihood and a minimax lower bound for the jointly tuned untied family. Further loop iterations improve the approximation bound without adding parameters, while increasing computation and the fitted-class complexity bound. For targets of known Hölder smoothness, looped residual feedforward networks and post-layer-normalized Transformers attain the minimax polynomial rate up to logarithmic factors with a fixed number of bounded real parameters. At sufficiently large fixed budgets, looped worst-case risk vanishes while optimal worst-case untied risk remains bounded away from zero. The loop-to-untied risk ratio also tends to zero under specified growing-budget conditions. Regression, binary response, and energy-based generative models illustrate the theory.
Comments79 pages, 5 figures; includes supplementary appendices