arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.00991cs.LG

适配器丛:分割RLVR预算优于集中使用

Adapter Thickets: Splitting an RLVR Budget Beats Concentrating It

Jonathan Williams, Esin Tureci, Karthik R. Narasimhan

首次发表
浏览论文内容

中文总结 AI 辅助

本研究证明RLVR预算集中训练单个适配器会因错误相关性增加而损害多数投票,提出将预算分割为多个适配器丛,在多种设置下显著提升投票准确率。

中文摘要 AI 辅助

对采样完成结果进行多数投票是测试时扩展的主力方法,而基于可验证奖励的强化学习(RLVR)则是提升每次完成质量的主力方法。标准流程将两者组合:先用RLVR训练一个策略,然后多次采样并投票。我们证明这种组合是有损的。投票只能纠正投票者未共有的错误,而RLVR使策略变得尖锐,导致其样本越来越倾向于犯相同的错误。在每种方法每个问题恰好抽取160个完成结果的条件下,在完整RLVR预算上训练单个LoRA适配器,在我们测试的每个模型(1.5B-8B)上均提升了单样本准确率。然而,在四个模型中的三个上,多数投票准确率低于未训练基础模型,最多低4.8个百分点。损害在训练过程中累积:投票者的错误相关性稳步增加,多数投票准确率在早期达到峰值后下降最多7.0个百分点。原因是集中,而非RLVR本身。我们将相同的数据和训练预算分割到K个LoRA适配器上,每个适配器在各自随机不相交的数据分片上训练,并将结果称为适配器丛。在所有16种(模型,K)设置中,适配器丛的投票结果优于完全训练的适配器,且对于K≥4,其准确率保持在基础模型的0.8个百分点以内或高于基础模型。在适配器丛成员的步数下提前停止的单个适配器,作为强对照,在小K情况下与适配器丛表现相当。对于K≥8,适配器丛保留了更多RLVR的单样本增益,并在八种设置中的六种中投票结果优于该对照。集中的代价也随投票数量增长:从16票到160票,适配器丛相对于完全训练适配器的领先优势从1.3个百分点扩大到3.3个百分点。当计划是采样并投票时,RLVR预算应更广泛而非更深入地使用。

英文摘要

Majority voting over sampled completions is the workhorse of test-time scaling, and reinforcement learning with verifiable rewards (RLVR) is the workhorse for making each completion better. The standard pipeline composes the two: train one policy with RLVR, then sample it many times and vote. We show that this composition is lossy. A vote can only overturn mistakes that its voters do not share, and RLVR sharpens a policy so that its samples increasingly make the same mistakes. With every method drawing exactly $160$ completions per problem, training a single LoRA adapter on the full RLVR budget raises single-sample accuracy on every model we test ($1.5$B-$8$B). Yet on three of four models it leaves the majority vote below that of the untrained base model, by up to $4.8$ points. The damage builds during training: voter errors grow steadily more correlated, and the majority vote accuracy peaks early before falling by up to $7.0$ points. The cause is concentration, not RLVR itself. We split the same data and training budget across $K$ LoRA adapters, each trained on its own random disjoint shard, and call the result an adapter thicket. Thickets out-vote the fully trained adapter in all $16$ (model, $K$) settings, and for $K{\geq}4$ they stay within $0.8$ points of the base model or above it. A single adapter stopped early, at a thicket member's step count, is a strong control that matches thickets for small $K$. For $K{\geq}8$, thickets keep more of RLVR's single-sample gain and out-vote this control in six of eight settings. The cost of concentration also grows with the number of votes: from $16$ to $160$ votes, the thicket's lead over the fully trained adapter widens from $1.3$ to $3.3$ points. When the plan is to sample and vote, an RLVR budget is better spent broad than deep.

发表机构

  • Princeton University(普林斯顿大学)

机构由 AI 辅助整理,请以论文原文为准。

↑