发表机构
Dartmouth College; IISc Bangalore; Xero(达特茅斯学院; 印度科学研究所班加罗尔分校; Xero)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出架构采样,一种无需训练的方法,通过重用解码器层块改变前向计算路径,在冻结视觉语言模型中生成多样候选,平均提升pass@9达6.58个百分点,并改善测试时强化学习效果。
AI 中文摘要
测试时扩展通常通过从冻结模型中采样多个响应来寻求更好的答案,然而传统的温度采样会沿着相同的固定计算路径生成每个候选。我们引入了架构采样,这是一种无需训练的方法,通过重用解码器层的选定块,经由不同的前向计算生成候选。改变块的位置和重复次数引入了计算多样性,而无需更新模型权重或添加辅助参数。在五个Qwen检查点和十二个多模态基准上,架构采样在相同的九个候选预算下,相较于标准路径温度采样,平均将pass@9提高了6.58个百分点。重用早期层带来了最大的增益,并且即使在贪婪解码下,候选覆盖率的提升也持续存在。生成的候选显示出较低的词汇重叠,并且在用作无标签测试时强化学习的rollout时提高了准确性。这些发现将我们的架构采样的益处扩展到候选覆盖率之外,展示了从模型自身输出中进行更有效学习的效果。
英文摘要
Test-time scaling often seeks better answers by sampling multiple responses from a frozen model, yet conventional temperature sampling generates every candidate along the same fixed computation path. We introduce architectural sampling, a training-free method that generates candidates through distinct forward computations by reusing selected blocks of decoder layers. Varying the block location and repetition count introduces computational diversity without updating model weights or adding auxiliary parameters. Across five Qwen checkpoints and twelve multimodal benchmarks, architectural sampling improves pass@9 over standard-path temperature sampling by 6.58 percentage points on average at the same nine-candidate budget. Reusing early layers yields the strongest gains, and the improvement in candidate coverage persists even under greedy decoding. The resulting candidates show lower lexical overlap and improve accuracy when used as rollouts for label-free test-time reinforcement learning. These findings extend the benefits of our architectural sampling beyond candidate coverage, demonstrating more effective learning from a model's own outputs.