模块化范数RandOpt:通过架构感知扰动实现群体高效集成
Modular Norm RandOpt: Population-Efficient Ensembling through Architecture-Aware Perturbations
浏览论文内容
中文总结 AI 辅助
针对RandOpt全局扰动忽略模块异质性的问题,提出模块化范数RandOpt,利用模块级自然范数采样,以更少候选在多个任务和模型规模上提升集成准确率。
中文摘要 AI 辅助
RandOpt对权重扰动的语言模型进行采样,并通过多数投票对排名靠前的候选模型进行集成,但其全局扰动尺度忽略了异质模块的几何结构。我们提出模块化范数RandOpt(Modular Norm RandOpt),一种架构感知的采样方法,使用模块级自然范数和校准尺度,同时保留选择和投票机制。它在Countdown任务上使用少3倍的候选模型,在GSM8K上至少少12倍的候选模型,即超越RandOpt,并相应节省了墙钟时间。在七个任务和三个Qwen规模(0.5B至3B)上的评估显示,在Countdown、GSM8K和MATH-500上,每个规模的平均准确率均高于RandOpt。这些增益扩展到Llama 3.2 3B和Gemma 3 4B在Countdown和GSM8K上的表现。在Qwen2.5-1.5B上,我们的集成在可比的主运行评估预算下,在两个任务上也实现了比迭代基线更高的平均准确率。在GSM8K上,尾部密度诊断表明仅减少1.2至1.8倍的候选数量,而大多数集成改进与更有利的正确专家支持相关。这些结果凸显了扰动几何作为围绕预训练模型进行群体高效、无梯度搜索的关键设计选择。
英文摘要
RandOpt samples weight-perturbed language models and ensembles top-ranked candidates through plurality voting, but its global perturbation scale ignores heterogeneous module geometry. We propose Modular Norm RandOpt, an architecture-aware sampling method using module-wise natural norms and calibrated scales while preserving selection and voting. It outperforms RandOpt using $3\times$ fewer candidates on Countdown and at least $12\times$ fewer on GSM8K, with corresponding wall-clock savings. Evaluations across seven tasks and three Qwen scales ($0.5$B--$3$B) show higher mean accuracy than RandOpt on Countdown, GSM8K, and MATH-500 at every scale. The gains extend to Llama 3.2 $3$B and Gemma 3 $4$B on Countdown and GSM8K. On Qwen2.5-1.5B, our ensembles also achieve higher mean accuracy than iterative baselines on both tasks at comparable main-run evaluation budgets. On GSM8K, a tail-density diagnostic implies only a $1.2$--$1.8\times$ candidate reduction, while most ensemble improvement is associated with more favorable correct-expert support. These results highlight perturbation geometry as a key design choice for population-efficient, gradient-free search around pretrained models.
发表机构
- The University of Osaka(大阪大学)
机构由 AI 辅助整理,请以论文原文为准。