arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HyperMC:面向随机梯度马尔可夫链蒙特卡洛的多保真度超参数调优

HyperMC: Multi-Fidelity Hyperparameter Tuning for Stochastic Gradient MCMC

Ming Tan, Xiyun Jiao

arXiv 2609.02138首次发表:更新:

发表机构

Southern University of Science and Technology(南方科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对SGMCMC超参数调优的难题,提出HyperMC多保真度调优框架及Robust HyperMC变体,在多个基准任务上实现了更优的后验近似或预测校准,且调优结果更稳定可复现。

AI 中文摘要

随机梯度马尔可夫链蒙特卡洛(SGMCMC)方法可实现可扩展的贝叶斯推断,但其性能强烈依赖于步长、小批量大小和蛙跳步数等超参数。由于大多数SGMCMC算法缺乏Metropolis-Hastings接受率,基于标准接受率的调优方法无法直接应用。我们提出HyperMC,一种多保真度调优框架,将Hyperband式资源分配与核斯坦差异(KSD)评估相结合。通过运行多个 successive-halving bracket,HyperMC在固定计算预算下平衡了对连续超参数空间的广泛探索与对有前景配置的日益准确评估。我们进一步引入Robust HyperMC,其采用全局网格初始化后接精英引导的局部优化,以降低对随机候选生成和有限预算下噪声评估的敏感性。在估计KSD的适当近似与集中条件下,我们证明 successive-halving 组件以高概率从采样候选中选出近最优配置,并推导了成功选择所需的充分计算预算。在逻辑回归、概率矩阵分解和贝叶斯神经网络上的实验表明,HyperMC相比MAMBA、网格搜索和启发式基线方法,在 posterior 近似或预测校准方面有所提升,而Robust HyperMC则能产生更稳定、可复现的调优结果。

英文摘要

Stochastic gradient Markov chain Monte Carlo (SGMCMC) methods enable scalable Bayesian inference, but their performance depends strongly on hyperparameters such as the step size, mini-batch size, and number of leapfrog steps. Since most SGMCMC algorithms lack a Metropolis-Hastings acceptance rate, standard acceptance-based tuning methods are not directly applicable. We propose HyperMC, a multi-fidelity tuning framework that combines Hyperband-style resource allocation with kernel Stein discrepancy (KSD) evaluation. By running multiple successive-halving brackets, HyperMC balances broad exploration of a continuous hyperparameter space with increasingly accurate evaluation of promising configurations under a fixed computational budget. We further introduce Robust HyperMC, which uses global grid initialization followed by elite-guided local refinement to reduce sensitivity to random candidate generation and noisy finite-budget evaluations. Under suitable approximation and concentration conditions for the estimated KSD, we establish that the successive-halving component selects a near-optimal configuration among the sampled candidates with high probability and derive a sufficient computational budget for successful selection. Experiments on logistic regression, probabilistic matrix factorization, and Bayesian neural networks show that HyperMC improves posterior approximation or predictive calibration relative to MAMBA, grid search, and heuristic baselines, while Robust HyperMC yields more stable and reproducible tuning results.

Comments56 pages, 14 figures, 9 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑