arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

结构化特征过拟合,随机特征涌现泛化

Structured Features Overfit Where Random Features Grok

Chon-Fai Kam, Miloud Bessafi, Frederic Cadet

arXiv 2609.15047首次发表:更新:

发表机构

University Paris City & University of Reunion; Dipartimento di Fisica e Chimica Emilio Segrè, Università degli Studi di Palermo; EnergyLab, University of Reunion; PEACCEL, AI for Biologics(巴黎城市大学与留尼汪大学; 巴勒莫大学埃米利奥·塞格雷物理与化学系; 留尼汪大学能源实验室; PEACCEL生物制药人工智能公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究对比结构化与随机特征映射,发现结构化傅里叶特征在岭回归中不出现涌现延迟,其性能退化由活跃模式数而非容量比决定,揭示特征几何对泛化的关键作用。

AI 中文摘要

Xu、Vardi和Safran(ICML 2026)证明了在非结构化的随机高斯特征映射上,过参数化岭回归会发生“涌现”(grok)现象,即记忆与泛化之间的延迟随权重衰减的$1/\lambda$增长。我们表明,在结构化特征映射上,同样的延迟不会出现。对于定义在$\mathbb{Z}_p^2$上的带限傅里叶特征映射,若目标为可表达类内的单字符目标,在固定正权重衰减下扩大频带,会使峰值留出准确率从$1.00$单调降至$0.07$,在整个扫描过程中不存在“先记忆后泛化”的阶段。这种退化并非插值效应。它在容量比$q/n = 0.638$时出现,远低于插值阈值,其成因与高于该阈值时出现的精确零空间不同。真正具有清晰边界的是活跃支撑集。保持名义维度固定,将频带掩蔽回$1089$个活跃模式,留出准确率恢复至$1.000$且各随机种子间方差为零,而完整的$4225$模式频带则崩溃至$0.185$。活跃模式的数量通过经验Gram矩阵的教师加权谱起作用,而非通过容量比,这使得该结果关乎特征几何,而非对双重下降的重新表述。

英文摘要

Xu, Vardi and Safran (ICML 2026) prove that over-parameterized ridge regression over an unstructured random Gaussian feature map groks, with the delay between memorization and generalization growing as $1/λ$ in the weight decay. We show that on a structured feature map the same delay does not appear. For a band-limited Fourier feature map over $\mathbb{Z}_p^2$ carrying a single-character target that lies inside the expressible class, enlarging the band at fixed positive weight decay drives peak held-out accuracy monotonically from $1.00$ to $0.07$, with no memorize-then-generalize regime anywhere along the sweep. The degradation is not an interpolation effect. It sets in at capacity ratio $q/n = 0.638$, far below the interpolation threshold, on separate grounds from the exact null space that appears above it. What does have a sharp boundary is the active support. Holding the nominal dimension fixed and masking the band back to $1089$ active modes restores held-out accuracy of $1.000$ with zero variance across seeds, while the full $4225$-mode band collapses to $0.185$. The number of active modes acts through the teacher-weighted spectrum of the empirical Gram matrix and not through the capacity ratio, which makes this a statement about feature geometry and not a restatement of double descent.

Comments13 pages, 2 figures, 3 tables. Submitted to OPT 2026 (18th Annual Workshop on Optimization for Machine Learning)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑