arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语言模型微调中基于奖励感知的进化策略种群规模缩放

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning

Sung Cho, Gyubin Han

arXiv 2607.19408首次发表:更新:

AI 中文总结

研究在大语言模型微调中,进化策略种群规模缩放与奖励设计及归一化的关系。通过实验发现交叉熵与二元奖励微调种群规模结论差异大,主要因奖励设计和归一化,禁用归一化可改进小种群模型表现,揭示小种群失败可能是实现问题而非内在限制。

AI 中文摘要

使用进化策略(ES)微调大语言模型很有吸引力,因为它内存高效、可并行且与黑盒或离散奖励兼容。但其种群规模结论差异很大:交叉熵(CE)奖励微调时N = 1成功,而二元奖励训练通常需要N≈30。我们表明这种差距主要在于奖励设计和归一化,而非种群规模。在我们研究的能力模型体系中,z分数优势归一化会导致N = 2失败。禁用归一化后,N = 2的二元奖励ES在0.5B - 7B的能力模型上对GSM8K和TREC有改进,而归一化变体则崩溃或退化。这种小N风险由奖励粒度决定,二元准确率奖励会产生零优势概率q,它取决于基础准确率、批量大小和对内正确性相关性。在Qwen2.5 - Instruct/GSM8K上的零训练探针在12种配置下平均绝对误差为0.020,与公式匹配,并发现此能力模型体系中可用性阈值Navail很小。这意味着并非N = 2普遍足够,而是能力模型二元ES中的小种群失败可能是实现工件而非内在种群限制。

英文摘要

Using Evolutionary Strategies (ES) for fine-tuning large language models is attractive because it is memory-efficient, parallel, and compatible with black-box or discrete rewards. Yet its population-size conclusions conflict sharply: fine-tuning with cross-entropy (CE) reward succeeds with $N=1$, while binary-reward training often needs $N \approx 30$. We show this gap is largely about reward design and normalization, not population size. In the capable-model regime we study, z-score advantage normalization can cause $N=2$ to fail. Disabling normalization lets binary-reward ES with $N=2$ improve on GSM8K and TREC across capable models spanning 0.5B-7B, where the normalized variant collapses or degrades. This small-$N$ risk is set by reward granularity: binary accuracy reward induces a zero-advantage probability $q$ that depends in closed form on base accuracy, batch size, and intra-pair correctness correlation; a zero-training probe on Qwen2.5-Instruct/GSM8K matches the formula with mean absolute error 0.020 across 12 configurations and finds the availability threshold $N_{\mathrm{avail}}$ to be small in this capable-model regime. The implication is not that $N=2$ is universally sufficient, but that small-population failure in capable-model binary ES can be an implementation artifact rather than an intrinsic population limit.

Comments15 pages, 2 figures. Accepted at the 4th HiLD (High-dimensional Learning Dynamics) Workshop, ICML 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑