arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

对称感知特征学习:多指标模型的多项式分离

Symmetry-Aware Feature Learning: A Polynomial Separation for Multi-Index Models

Jivan Waber, Vanessa Piccolo, Yatin Dandi, Florent Krzakala

arXiv 2610.08420首次发表:更新:

发表机构

Information, Learning and Physics Laboratory; Statistical Physics of Computation Laboratory; École Polytechnique Fédérale de Lausanne (EPFL)(信息、学习与物理实验室; 统计物理与计算实验室; 洛桑联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文证明对称感知学习(权重共享或全群数据增强)相比对称无关学习,在多指标模型上实现多项式样本复杂度优势,并揭示两阶段机制。

AI 中文摘要

我们建立了对称感知与对称无关特征学习之间的多项式样本复杂度分离。我们研究了高维高斯协变量在$\mathbb{R}^d$中的增长秩多指标模型,其中$r=\Theta(d^\delta)$个教师方向形成循环对称轨道,且$0<\delta<1/2$。我们比较了利用这种结构的三种方式:架构权重共享、全对称群上的数据增强,以及不利用对称性的学习。具体而言,我们分析了对称绑定的卷积网络、非绑定网络,以及使用全群数据增强训练的相同非绑定网络,采用带有相关性损失的球面在线SGD。对于信息指数$p\ge3$的一类多项式链接,我们证明了匹配的样本复杂度界限(对数因子内):绑定和增强学习者在$\widetilde{\Theta}(d^{p-1})$个样本内实现弱方向恢复,而对称无关学习者需要$\widetilde{\Theta}(rd^{p-1})$个样本。对于纯二次Hermite链接,相同的分离在教师子空间的弱恢复中成立,样本复杂度分别为$\widetilde{\Theta}(d)$和$\widetilde{\Theta}(rd)$。因此,全群数据增强匹配了架构权重共享的样本效率,并且两者相比不利用对称性的训练提供了多项式优势。对于$p\ge3$,证明揭示了一个两阶段机制:初始化时的波动选择教师轨道中的一个方向,随后局部增长将其重叠放大到弱恢复尺度,而竞争重叠保持在其初始化尺度附近。

英文摘要

We establish a polynomial sample complexity separation between symmetry-aware and symmetry-agnostic feature learning. We study growing-rank multi-index models with high-dimensional Gaussian covariates in $\mathbb{R}^d$ and $r=Θ(d^δ)$ teacher directions forming a cyclic symmetry orbit, where $0<δ<1/2$. We compare three ways of exploiting this structure: architectural weight sharing, data augmentation over the full symmetry group, and learning without access to the symmetry. In particular, we analyze a symmetry-tied convolutional network, an untied network, and the same untied network trained with full-group data augmentation, using spherical online SGD with correlation loss. For a class of polynomial links with information exponent $p\ge3$, we prove matching sample complexity bounds up to logarithmic factors: the tied and augmented learners achieve weak directional recovery in $\widetildeΘ(d^{p-1})$ samples, whereas the symmetry-agnostic learner requires $\widetildeΘ(rd^{p-1})$. For the pure quadratic Hermite link, the same separation holds for weak recovery of the teacher subspace, with sample complexities $\widetildeΘ(d)$ and $\widetildeΘ(rd)$, respectively. Thus, full-group data augmentation matches the sample efficiency of architectural weight sharing, and both provide a polynomial advantage over training without symmetry. For $p\ge3$, the proof reveals a two-stage mechanism: fluctuations at initialization select one direction in the teacher orbit, after which localized growth amplifies its overlap to the weak recovery scale while competing overlaps remain near their initialization scale.

Comments71 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑