发表机构
University College London; Institut Polytechnique de Paris; University of Oxford(伦敦大学学院; 巴黎理工学院; 牛津大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文探究在适度评估预算下,贝叶斯优化能否在神经丛中高效找到强单专家,提出在权重空间随机线性嵌入中应用贝叶斯优化的无梯度方法,在Qwen2.5-Instruct模型基准测试中,该方法评估成本更低且性能优于随机优化。
AI 中文摘要
无梯度后训练已成为大语言模型(LLMs)梯度优化的极具吸引力的替代方案,但现有方法成本仍较高。本文研究在适度评估预算下,结构化搜索能否识别强单专家。受有用权重更新位于低维子空间的证据启发,我们在权重空间的随机线性嵌入中应用贝叶斯优化。该方法无需反向传播,使用高斯过程代理高效引导候选评估。在参数规模为0.5B至3B的Qwen2.5-Instruct模型的多个推理基准测试中,使用少五倍候选评估的贝叶斯优化,性能与随机优化(RandOpt)相当或更优。这些结果表明,代理引导搜索可大幅降低无梯度后训练的评估成本,同时生成更强的可部署单专家。
英文摘要
Gradient-free post-training has emerged as a compelling alternative to gradient-based optimization for large language models (LLMs), but existing approaches remain costly. We ask whether structured search can identify a strong single expert under a modest evaluation budget. Motivated by evidence that useful weight updates lie in low-dimensional subspaces, we apply Bayesian optimization within a random linear embedding of weight space. Our method requires no backpropagation and uses a Gaussian process surrogate to guide candidate evaluations efficiently. Across several reasoning benchmarks with Qwen2.5-Instruct models from 0.5B to 3B parameters, Bayesian optimization using five times less candidate evaluations matches or exceeds RandOpt. These results show that surrogate-guided search can substantially reduce the evaluation cost of gradient-free post-training while producing stronger deployable single experts.