稀疏数据增强的优化及其可证明保证
Sparse Data Augmentation for Optimization with Provable Guarantees
浏览论文内容
中文总结 AI 辅助
针对几何机器学习中数据增强计算昂贵的问题,提出用固定稀疏变换样本近似完全增强,证明梯度下降在更少查询下达到ε-稳定点,优于完全增强GD和group-SGD。
中文摘要 AI 辅助
在几何机器学习中出现的非凸优化问题里,数据增强通常被用来通过对数据的变换进行经验损失平均来促进不变性。然而,计算完全增强的目标函数需要访问变换群$G$的每一个元素,当$G$很大或只能通过采样访问时,这可能会极其昂贵。我们研究是否可以用一个在优化前获取并在之后重复使用的小型固定变换样本来近似完全增强。在适当的正则性条件下,我们证明,以至少$1-\delta$的概率,在由此产生的稀疏增强目标上运行的梯度下降(GD)使用$\mathcal{O}\bigl((\log |G|+\log(1/\delta))/\varepsilon^2\bigr)$次群变换预言机查询,返回完全增强目标的一个$\varepsilon$-稳定点。相比之下,标准的群随机梯度下降(group-SGD)在每次迭代时采样一个新的变换,使用$\mathcal{O}(1/\varepsilon^4)$次变换查询。因此,使用固定稀疏增强的梯度下降所需的变换查询次数少于应用于完全增强目标的GD和group-SGD。我们的证明技术可能具有独立的意义,通过群诱导算子的谱性质和表示论工具,建立了全群平均梯度场被随机群平均均匀逼近的结果。
英文摘要
In nonconvex optimization problems arising in geometric machine learning, data augmentation is commonly used to promote invariance by averaging empirical losses over transformations of the data. Computing the fully augmented objective, however, requires access to every element of the transformation group $G$, which may be prohibitively expensive when $G$ is large or accessible only through sampling. We study whether full augmentation can instead be approximated using a small, fixed sample of transformations acquired before optimization and reused thereafter. Under suitable regularity conditions, we show that, with probability at least $1-δ$, gradient descent (GD) on the resulting sparsely augmented objective returns an $\varepsilon$-stationary point of the fully augmented objective using $\mathcal{O}\bigl((\log |G|+\log(1/δ))/\varepsilon^2\bigr)$ group-transformation-oracle queries. By comparison, standard group stochastic gradient descent (group-SGD), which samples a fresh transformation at every iteration, uses $\mathcal{O}(1/\varepsilon^4)$ transformation queries. Therefore, gradient descent with fixed sparse augmentation requires fewer transformation queries than both GD applied to the fully augmented objective and group-SGD. Our proof techniques, which may be of independent interest, establish a uniform approximation of the full group-averaged gradient field by a random group average using spectral properties of group-induced operators and tools from representation theory.
发表机构
- Harvard John A. Paulson School of Engineering and Applied Sciences, Harvard University(哈佛大学约翰·A·保尔森工程与应用科学学院)
机构由 AI 辅助整理,请以论文原文为准。