arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

随机特征的自适应正则化:具备神谕速率保证的邻域早停规则

Adaptive Regularization for Random Features: A Neighboring Early-Stopping Rule with Oracle-Rate Guarantees

Caixing Wang, Zhibo Chen, Yue Wang

arXiv 2608.25513首次发表:更新:

发表机构

Southeast University; University of Science and Technology of China(东南大学; 中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对带随机特征的核岭回归提出邻域早停规则,无需预先知晓未知参数即可选正则化参数,可达到神谕多项式学习速率,经实验验证了其性能。

AI 中文摘要

随机特征方法为核岭回归(KRR)提供了可扩展的近似,但能达到神谕学习速率的正则化参数依赖于未知的光滑性和容量参数。在本研究中,我们为采用随机特征的核岭回归(KRR-RF)提出了一种用于自适应正则化的邻域早停规则。该方法使用逆正则化上均匀分布的网格,仅比较相邻估计量,相较于标准的全对Lepskii型方法,减少了差异比较的数量。邻域差异及其经验复杂度项均可直接在随机特征空间中计算,无需构建精确的核Gram矩阵。我们为邻域KRR-RF估计量建立了高概率比较界,并证明在标准源条件、容量条件,以及合适的网格和随机特征预算条件下,所选估计量可达到神谕多项式学习速率(仅差对数因子)。该结果使得正则化参数的选择无需预先知晓源指数和容量指数,且涵盖了模型完全指定和部分 misspecified(误指定)两种情况。我们的分析基于经验随机特征有效维度,该维度将可观测的停止阈值与随机特征模型的总体复杂度关联起来。仿真和真实数据实验展示了所提方法相较于标准调优过程的预测性能和计算特性。

英文摘要

Random feature methods provide a scalable approximation to kernel ridge regression (KRR), but the regularization parameter that yields the oracle learning rate depends on unknown smoothness and capacity parameters. In this work, we propose a neighboring early-stopping rule for adaptive regularization in KRR with random features (KRR-RF). The method uses a grid that is uniform in inverse regularization and compares only adjacent estimators, reducing the number of discrepancy comparisons relative to standard all-pairs Lepskii-type procedures. Both the neighboring discrepancy and its empirical complexity term can be computed directly in the random feature space, without constructing the exact kernel Gram matrix. We establish a high-probability comparison bound for neighboring KRR-RF estimators and show that, under standard source and capacity conditions together with suitable grid and random feature budget conditions, the selected estimator attains the oracle polynomial learning rate up to logarithmic factors. The result allows the regularization parameter to be selected without prior knowledge of the source and capacity exponents and covers both well-specified and partially misspecified regimes. Our analysis is based on an empirical random feature effective dimension that connects the observable stopping threshold with the population complexity of the random feature model. Simulation and real-data experiments illustrate the prediction performance and computational behavior of the proposed method in comparison with standard tuning procedures.

Comments31 pages, 10 figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑